keyphrase GitHub

How it works

The same input always returns the same phrases. Each phrase is a slice of the string you passed. StartByte and EndByte are indexes into that string.

Extraction pipeline from the caller's text to the selected phrases 1 Segment Unicode sentence and word boundaries keep the original byte offsets. Abbreviations, initials, and forms such as "U.S." stay one sentence. 2 Fix tokens Hyphenated compounds stay one token. German compounds stay whole. Elided articles (l', d', dell') and English 's split off. 3 Tag A model labels every token: noun, name, verb, adverb, and the rest. Verbs and confident adverbs never appear inside a phrase. 4 Candidates A run is at most three words, and it stays inside one sentence. Connectors such as of, de, and für sit inside, never at an end. 5 Score Equal-weight mean of six features, each between 0 and 1. frequency early spread capitalization cohesion rarity 6 Select Drop a shorter phrase, a near duplicate, or an overlap. Keep the earliest free span, then scale scores to the best one.
  1. Segment. Unicode sentence and word boundaries (UAX #29) keep the original byte offsets. Known abbreviations, initials, and dotted forms such as "U.S." stay in one sentence.
  2. Fix tokens. Hyphenated compounds such as "e-mail" stay one token. Elided articles ("l'", "d'", "dell'") and English "'s" split off, with straight or curly apostrophes. German compounds stay single tokens.
  3. Tag. A small model per language labels each token. Verbs and helper verbs never start, end, or appear inside a phrase. Confident adverbs are treated the same way. A word the tagger rarely calls a noun or a name cannot stand as the head of a phrase.
  4. Candidates. Contiguous runs of at most three words inside one sentence, counting connectors such as "of", "de", and "für". Punctuation, line breaks, verbs, and adverbs end a run. A connector can sit inside a phrase, never at either end, and never more than two in a row. A single word is kept only when it is a name, capitalized like a name, or a common noun that is rare in general text.
  5. Score. The score is the equal-weight mean of how often the phrase occurs, how early it appears, how many sentences it spans, how often it looks like a name, how tightly its words stay together, and how uncommon its content words are in a general corpus.
  6. Select. In rank order, skip a phrase contained in a longer one that ranks at least as high, skip near duplicates, skip anything that overlaps a phrase already chosen, and keep the earliest non-overlapping occurrence. Scores are then divided by the best selected score.

Tagger

The tagger is a left-to-right averaged perceptron. For each word it scores labels from the word, its shape, its neighbours, and the labels it has just given. Frequent words that almost always have one label are looked up instead. The weights were learned from Universal Dependencies treebanks, split with this extractor's own tokenizer.

On each treebank's held-out sentences, verb precision is the share of words labelled as verbs that really are verbs. Verb recall is the share of real verbs that get the verb label.

LanguageTraining dataWords correctVerb precisionVerb recall
EnglishEWT, 209k tokens94.0%96.6%97.4%
FrenchGSD, 346k tokens96.8%96.7%98.1%
GermanGSD, 260k tokens93.9%96.7%97.4%
SpanishGSD, 376k tokens95.6%97.9%96.9%
ItalianParlaMint, PUD, TWITTIRO, MarkIT, 90k tokens94.4%95.5%96.4%

Italian uses only openly licensed treebanks, about a quarter as much data as the others, so it is the weakest of the five. A model is decoded the first time its language is used. A language with no model still runs, and it keeps verbs.

Limitations

  • You pass the language code on every call.
  • The input should already be article text. Menus, footers, and other page chrome can come back as phrases.
  • The tagger misses some verbs, and sometimes drops a word that is not a verb.
  • Names longer than three words are split.
  • A score is a ranking signal for one document. The best phrase in that document scores 1. Scores from two documents are not comparable.

Accuracy is measured on texts that are not in the repository. The tools that score those texts live with the source, and they run on any text you have locally.