Language coverage by NLP service
The Bitext NLP matrix maps each linguistic service to the languages where it is available: language identification, sentence segmentation, tokenization, word segmentation, lemmatization, decompounding, spelling, POS tagging, entity extraction, phrase extraction and parsing.
20 language variants
11 NLP service
Decompounding included
Samples available
Base languages
Language matrix coverage
Variants
Language variants in the matrix summary
Max total resources
77 languages + 20 variants for broad services
NLP services
From language ID to parsing
What each service covers
Each card shows base-language coverage, variant coverage and total resources for that service.
Language ID
detect language in mixed text
Sentence Seg.
split text into sentences
Tokenization
split sentences into words
Word Seg.
no-space tokenization
Lemmatization
return roots for word forms
Decompounding
split compound words
Spelling
check spelling correctness
POS tagging
disambiguate parts of speech
Entities
detect names and special text
Phrase Extraction
noun, verb and prep phrases
Parsing
hierarchical sentence trees
Full pipeline
combine lexical, syntactic and semantic services
Exact coverage by NLP service
Language identification
20 variants
97 total
Detect language used in each sentence of a longer input text
Sentence segmentation
20 variants
97 total
Split text into sentences using language-specific punctuation rules
Tokenization
19 variants
92 total
Split sentences into words using language-specific space and punctuation rules
Word segmentation
2 variants
6 total
Split no-space text into words for languages such as Chinese, Japanese, Vietnamese and Thai
Lemmatization
20 variants
97 total
Return possible roots for word forms
Decompounding
6 variants
19 total
Split compound words into component words for compound-heavy languages
Spelling
20 variants
97 total
Check if a word is spelled correctly
POS tagging
15 variants
45 total
Disambiguate and return the part of speech for each word in a sentence
Entity extraction
19 variants
33 total
Detect proper names, places, organizations and special text such as phones or URLs
Phrase extraction
12 variants
20 total
Return constituents such as noun phrases, verb phrases and prepositional phrases
Parsing
13 variants
34 total
Produce a hierarchical sentence tree of words, phrases and clauses
Lexical, syntactic and semantic coverage
The matrix separates the services that prepare raw text at word level from the services that add grammar, structure and meaning.
Prepare raw text before it reaches downstream systems
These services normalize, split and validate text at the language and word level
Add grammar, entities and sentence-level structure
These services create richer linguistic signals for extraction, analysis and AI workflows
77 base languages
Broad lexical services such as language identification, sentence segmentation, lemmatization and spelling are available across this base list.
Regional coverage for market-specific behavior
Variants matter when vocabulary, spelling, morphology, entities or usage differ across markets
Proof buyers can open
These sample files and language specification PDFs show real lexical entries, lemmas, POS, morphology, entity information, frequency and language-specific attributes.
Portuguese
Download the full NLP service matrix
Download the updated Appendix C matrix to review Bitext language coverage by NLP service, including lexical normalization, decompounding, POS tagging, entity extraction, phrase extraction and parsing.
MADRID, SPAIN
Camino de las Huertas, 20, 28223 Pozuelo
Madrid, Spain
SAN FRANCISCO, USA
541 Jefferson Ave Ste 100, Redwood City
CA 94063, USA