Showing posts with label search engine. Show all posts
Showing posts with label search engine. Show all posts

Monday, May 14, 2007

The Future of Search Engines (outgoing links in Spanish)

On May 10, 2007 in the beautiful city of Santiago de Compostela I participated in a round table discussion on the future of search engines. The event was organized by FESABID: the tenth Spanish Days of Documentation. This is the main Spanish event for libraries and other institutions that manage very many documents. About 800 persons attended. The panel was formed by (in parenthesis appear the topics of the talk):

  • Richard Benjamins, iSOCO (Vertical search engines and Semantic Technology)
  • Eva Méndez, Universidad Carlos III (Dublin Core)
  • José Ramón Pérez Agüera, Universidad Complutense de Madrid (IR algorithms, Semantics and Natural Language Processing).
  • Antonio Pareja Lora, UCM, colaborador del Grupo de Ingeniería Ontológica (OEG) de la UPM (Ontologies)
  • Antonio S. Valderrábanos, Bitext.com (Natural Language Processing)
  • Isidro Aguillo, CINDOC/CSIC (moderator)

    The presentations given by the panel members formed an interesting mix of different by complementary views on the matter of search engines: an excellent starting point for setting up a new breakthrough project to advance the state of the art.

    The presentations of the whole conference are made available here. My presentation can be downloaded here.
  • Tuesday, April 10, 2007

    A New Generation Search Engines

    The Financial Times recently published a short note: “New tools to vie with Google”, which briefly describes –from a user point of view- some of the new generation search engines. It includes new search engines like:

    All those search engines allow queries in natural language, and they claim to use NLP and semantics to “understand” the content.

    In my understanding, there are two ways where NLP and Semantics can play a role:

    • Interpreting the user query. This is what most new search engines claim to do.
    • “Understanding” the content to be indexed. This requires that at index time, not only the individual words of the content (documents) are indexed, but indexing also considers NLP and Semantics. It is unclear to what extent those “new” search engines apply this for indexing. Yet another possibility is to launch the query against structured information (e.g. RDF).
    I usually summarize the above two points in respectively Semantics in Access and Semantics in the Source. The latter is obviously much harder than the former, and a real challenge for the next generation search engines.

    The short note of Paul Taylor also mentions several “search” engines for comparing prices of products available on the web, including

    • Shopzilla
    • Pricegrabber
    • Shopping.com

    While the note talks about those “shopping bots” for consumers, also providers (the businesses) may have interest in such tools for tracking their products at resellers’ sites and for automatically tracking their competitors’ products.

    Friday, April 06, 2007

    Searching content of images

    A few days ago I mentioned the game approach for annotating images of Louis von Ahn, that makes use of free computation cycles of humans. A straigthforward use of the side effect of playing the games milliones of times, is a search engine for images. Here you can try prototypes for two search image engines:

    1. Search for images according to overall subject of image
    2. Search for objects within images

    Although the amount of images considered is still limited (tens of thousands), the approach seems promising.