Articulo de referencia

Lesk algorithm

The Lesk algorithm is a classical algorithm for word sense disambiguation introduced by Michael E. Lesk in 1986. [ 1 ] It operates on the premise that words within a given conte...

The Lesk algorithm is a classical algorithm for word sense disambiguation introduced by Michael E. Lesk in 1986.[1] It operates on the premise that words within a given context are likely to share a common meaning. This algorithm compares the dictionary definitions of an ambiguous word with the words in its surrounding context to determine the most appropriate sense. Variations, such as the Simplified Lesk algorithm, have demonstrated improved precision and efficiency. However, the Lesk algorithm has faced criticism for its sensitivity to definition wording and its reliance on brief glosses. Researchers have sought to enhance its accuracy by incorporating additional resources like thesauruses and syntactic models.

Overview

The Lesk algorithm is based on the assumption that words in a given "neighborhood" (section of text) will tend to share a common topic. A simplified version of the Lesk algorithm is to compare the dictionary definition of an ambiguous word with the terms contained in its neighborhood. Versions have been adapted to use WordNet.[2] An implementation might look like this:

  1. for every sense of the word being disambiguated one should count the number of words that are in both the neighborhood of that word and in the dictionary definition of that sense
  2. the sense that is to be chosen is the sense that has the largest number of this count.

A frequently used example illustrating this algorithm is for the context "pine cone". The following dictionary definitions are used:

PINE 1. kinds of evergreen tree with needle-shaped leaves 2. waste away through sorrow or illness
CONE 1. solid body which narrows to a point 2. something of this shape whether solid or hollow 3. fruit of certain evergreen trees

As can be seen, the best intersection is Pine #1 ⋂ Cone #3 = 2.

Simplified Lesk algorithm

In Simplified Lesk algorithm,[3] the correct meaning of each word in a given context is determined individually by locating the sense that overlaps the most between its dictionary definition and the given context. Rather than simultaneously determining the meanings of all words in a given context, this approach tackles each word individually, independent of the meaning of the other words occurring in the same context.

"Una evaluación comparativa realizada por Vasilescu et al. (2004) [ 4 ] ha demostrado que el algoritmo Lesk simplificado puede superar significativamente la definición original del algoritmo, tanto en términos de precisión como de eficiencia. Al evaluar los algoritmos de desambiguación en los datos de todas las palabras en inglés de Senseval-2, miden una precisión del 58% utilizando el algoritmo Lesk simplificado en comparación con solo el 42% bajo el algoritmo original.

Nota: La implementación de Vasilescu et al. considera una estrategia de retroceso para las palabras no cubiertas por el algoritmo, consistente en el sentido más frecuente definido en WordNet. Esto significa que a las palabras cuyos posibles significados no generan ninguna superposición con el contexto actual o con otras definiciones de palabras se les asigna por defecto el sentido número uno en WordNet. [ 5 ]

Algoritmo LESK simplificado con sentido de palabra predeterminado inteligente (Vasilescu et al., 2004) [ 6 ]

La función COMPUTEOVERLAP devuelve el número de palabras en común entre dos conjuntos, sin tener en cuenta las palabras funcionales ni otras palabras de la lista de palabras vacías. El algoritmo original de Lesk define el contexto de una manera más compleja.

Críticas

Lamentablemente, el método de Lesk es muy sensible a la redacción exacta de las definiciones, por lo que la ausencia de una palabra puede alterar radicalmente los resultados. Además, el algoritmo solo detecta coincidencias entre las glosas de los sentidos considerados. Esta es una limitación importante, ya que las glosas de los diccionarios suelen ser bastante breves y no proporcionan suficiente vocabulario para establecer distinciones semánticas precisas.

Han aparecido numerosos trabajos que ofrecen diferentes modificaciones de este algoritmo. Estos trabajos utilizan otros recursos para el análisis (tesauros, diccionarios de sinónimos o modelos morfológicos y sintácticos): por ejemplo, pueden utilizar información como sinónimos, diferentes derivados o palabras de definiciones de palabras de definiciones. [ 7 ]

Variantes de Lesk

  • Lesk original (Lesk, 1986)
  • Adapted/Extended Lesk (Banerjee and Pederson, 2002/2003): In the adaptive lesk algorithm, a word vector is created corresponds to every content word in the wordnet gloss. Concatenating glosses of related concepts in WordNet can be used to augment this vector. The vector contains the co-occurrence counts of words co-occurring with w in a large corpus. Adding all the word vectors for all the content words in its gloss creates the Gloss vector g for a concept. Relatedness is determined by comparing the gloss vector using the Cosine similarity measure.[8]

There are a lot of studies concerning Lesk and its extensions:[9]

  • Wilks and Stevenson, 1998, 1999;
  • Mahesh et al., 1997;
  • Cowie et al., 1992;
  • Yarowsky, 1992;
  • Pook and Catlett, 1988;
  • Kilgarriff and Rosensweig, 2000;
  • Kwong, 2001;
  • Nastase and Szpakowicz, 2001;
  • Gelbukh and Sidorov, 2004.

See also

References

  1. Lesk, M. (1986). Automatic sense disambiguation using machine readable dictionaries: how to tell a pine cone from an ice cream cone. In SIGDOC '86: Proceedings of the 5th annual international conference on Systems documentation, pages 24-26, New York, NY, USA. ACM.
  2. Satanjeev Banerjee and Ted Pedersen. An Adapted Lesk Algorithm for Word Sense Disambiguation Using WordNet, Lecture Notes in Computer Science; Vol. 2276, Pages: 136 - 145, 2002. ISBN 3-540-43219-1
  3. Kilgarriff and J. Rosenzweig. 2000. English SENSEVAL:Report and Results. In Proceedings of the 2nd International Conference on Language Resourcesand Evaluation, LREC, Athens, Greece.
  4. Florentina Vasilescu, Philippe Langlais, and Guy Lapalme. 2004. Evaluating Variants of the Lesk Approach for Disambiguating Words. LREC, Portugal.
  5. Agirre, Eneko & Philip Edmonds (eds.). 2006. Word Sense Disambiguation: Algorithms and Applications. Dordrecht: Springer. www.wsdbook.org
  6. Florentina Vasilescu, Philippe Langlais, and Guy Lapalme. 2004. Evaluating Variants of the Lesk Approach for Disambiguating Words. LREC, Portugal.
  7. Alexander Gelbukh, Grigori Sidorov. Automatic resolution of ambiguity of word senses in dictionary definitions (in Russian). J. Nauchno-Tehnicheskaya Informaciya (NTI), ISSN 0548-0027, ser. 2, N 3, 2004, pp. 10–15.
  8. Banerjee, Satanjeev; Pedersen, Ted (2002-02-17). "An Adapted Lesk Algorithm for Word Sense Disambiguation Using WordNet". Computational Linguistics and Intelligent Text Processing. Lecture Notes in Computer Science. Vol. 2276. Springer, Berlin, Heidelberg. pp. 136–145. CiteSeerX 10.1.1.118.8359. doi:10.1007/3-540-45715-1_11. ISBN 978-3540457152.
  9. Roberto Navigli. Desambiguación del sentido de las palabras: una revisión , ACM Computing Surveys, 41(2), 2009, pp. 1–69.