Aperture - Java framework for getting data and metadata
Aperture is a Java framework for extracting and querying full-text content and metadata from various information systems. It could crawl and extract information from File system, Websites, Mail boxes and Mail servers. It supports various file formats like Office, PDF, Zip and lot more. Metadata information is extracted from image files. Aperture has a strong focus on semantics, metadata extracted could be mapped to predefined properties.
comments powered by Disqus
OpenPipe is an open source scalable platform for manipulating a stream of documents. A pipeline is an ordered set of steps / operations performed on a document to convert from its raw form to something ready to be put into the index. The operations performed on documents include language detection, field manipulation, POS tagging, entity extraction or submitting the document to a search engine.
SMILA is an extensible framework for building search solutions to access unstructured information in the enterprise. Besides providing essential infrastructure components and services, SMILA also delivers ready-to-use add-on components, like connectors to most relevant data sources. Using the framework as their basis will enable developers to concentrate on the creation of higher value solutions, like semantic driven applications etc.
GATE excels at text analysis of all shapes and sizes. It provides support for diverse language processing tasks such as parsers, morphology, tagging, Information Retrieval tools, Information Extraction components for various languages, and many others. It provides support to measure, evaluate, model and persist the data structure. It could analyze text or speech. It has built-in support for machine learning and also adds support for different implementation of machine learning via plugin.
UIMA analyzes large volumes of unstructured information in order to discover knowledge that is relevant to an end user. It is a framework with different set of components. The components include Language Identification, Language specific segmentation, Sentence boundary detection, Entity detection (person/place names) etc. The framework manages these components and the data flows between them.
Uima-connectors - uima connectors, solutions to build the bridge between some markup languages and t
OverviewUIMA-connectors aims mainly at offering solutions to build the bridge between some markup languages and the UIMA structure data, namely the CAS. In comparison, the Tika project aims at detecting and extracting metadata and structured text content from various type MIME documents. UIMA-connectors is more dedicated to perfom mapping from/to text formats to/from CAS, providing solutions for handling language formats such as eXtended Markup Language (XML), Comma Separated Value (CSV), whites
Welcome to the Open Text Livelink Connector Project! The Google Open Text Livelink Connector enables the Google Search Appliance to search and serve documents and other content stored in Open Text Livelink. Update: Livelink connector release 2.6.12 is now available. This is a patch release with some enhancements. This release replaces version 2.6.10. See the release notes for details on the changes in both versions. To install this release, use the 2.6.8 installer, and then update the Livelink c
Library providing dynamic connector capabilities to Google Web Toolkit (GWT) applications. Click to see demo. Click to see presentation. Questions? If you have any questions, please post them on http://groups.google.com/group/gwt-connectors and I will try to answer them. Using the library you will be able to: Draw connections between shapes. Change shape of the connection by dragging its sections. Keep shapes connected (glued) while you move them around. Add text to connections. Add decorations
Welcome to the Google Enterprise Connector Manager project! The Google Enterprise connector framework enables the Google Search Appliance to search and serve documents stored in non-Web repositories, such as enterprise content management systems. An enterprise content management (ECM) system provides a central repository for large numbers of documents. The Connector Manager is the central part of the connector framework for the Google Search Appliance. The Connector Manager itself manages creati
Constellio Open Source Enterprise Search is based on Apache Solr and using Google Search Appliances connectors architecture, it allows, with a single click, to find all relevant content in your organization (Web, email, ECM, CRM etc.).
This is the Itemscript JSON Toolkit for standard Java and GWT Java. Main components: A cross-platform GWT & standard Java JSON library, with convenient classes, parsers, and utilities. A RESTful connector API for retrieval of data (JSON, text & small binary files) over a variety of protocols. A simple in-memory database with a RESTful interface, useful as a mock server for testing & development, and for managing application state. A validator for the Itemscript Schema JSON schema language. The J