keyphrase GitHub

Reference

Extract returns up to limit non-overlapping phrases. ExtractWithCounts is the same call plus a count summary.

kws, err := extractor.Extract(text, "en", 10)
kws, counts, err := extractor.ExtractWithCounts(text, "en", 10)

Keyword

Each result is a span of the caller's text. Text is text[StartByte:EndByte]. Offsets are zero-based, half-open UTF-8 byte indexes.

FieldMeaning
TextThe phrase, sliced from the input.
StartByteFirst byte of the span.
EndByteByte after the span.
ScoreRanking signal in [0, 1].

Score is relative to this document. The best phrase scores 1. It is not a probability, and scores are not comparable across documents.

Order

Results are ordered by score descending, then by earlier StartByte, then by Text. Spans never overlap. The returned text is always a slice of the input.

Limit

The default limit is 5. limit must be at least 1. There is no maximum: a limit greater than the number of phrases returns all of them. A limit below 1 returns ErrInvalidLimit.

Language

language is the code of a package the program imported: en, fr, de, es, or it. A code whose package was not imported returns ErrUnsupportedLanguage.

Empty text

Empty or whitespace-only text is success and returns no phrases.

Counts

ExtractWithCounts adds:

  • Words — words in the text. Punctuation is not counted. A hyphenated compound such as "e-mail" is one word.
  • Extracted — every candidate phrase.
  • Ranked — how many of those were returned.
  • ExtractedMoreThanOneWord and RankedMoreThanOneWord — the same counts for phrases of two or more words, including connectors.
  • MoreThanOneWord — every extracted phrase of two or more words, best-ranked first. Its length equals ExtractedMoreThanOneWord.

The Try page always uses this summary. A 64 KiB clip applies there, and to text loaded from a URL. The library call itself does not reject longer text.