Explore real Bitext language samples and specifications
Use this resource page to inspect actual Bitext sample files and language specification documents. The goal is not to explain the pipeline again, but to give technical buyers direct access to the linguistic evidence behind it.
Language specifications
Morphological attributes
Named entities
Technical PDFs
Download sample data and language specifications
Use the XLSX sample files to inspect data structure. Use the PDF specifications to understand the feature sets available for each language.
Kazakh
Specification includes inflectional, derivational and extended forms, named entities, frequency and offensive-language flags.
Armenian
Specification includes POS, lemma, tense-mood, person, number, case, definiteness, degree and named entities.
Slovak
Specification includes perfective, tense, person, number, gender, case, degree, negative and entity-type information.
Mongolian
Specification includes tense-mood, aspect, polarity, person, number, case, degree and reflexiveness.
Russian
Use the sample and specification documents to inspect Russian lexical and morphological resource structure.
Portuguese
Specification includes tense, mood, person, number, gender, named entities and Brazilian Portuguese regional variant data.
Malayalam
Specification includes affirmative, tense, mood, person, gender, number, case, formality and named entity information.
Urdu
Specification includes tense, mood, gender, number, person, case, possessive attributes and named entities.
Catalan
Specification includes tense, mood, person, number, gender, derivational forms and common compound words.
The technical signals inside the samples
The sample files and specifications show the linguistic features that matter when Bitext is used as infrastructure for search, RAG, entity extraction, document intelligence and multilingual AI.
Lemma and form
Canonical forms and surface forms used for normalization and matching
POS and morphology
Part of speech plus attributes such as tense, mood, person, number, gender and case
Named entities
Entity-type signals for names, places, companies and organizations
Frequency
Relative frequency information from representative language corpora
Offensive flag
Information indicating whether a form may be offensive in specific contexts
Language-specific phenomena
Clitics, compounds, postposition suffixes, regional variants and other language-specific details
Use these documents for deeper evaluation
The sample library is useful for language-specific inspection. These two documents give buyers the broader technical proof: complete NLP service coverage and detailed lexical resource structure.
Lexical Data Resources
Detailed reference for Bitext lexical data resources, morphological features, frequency data, named entities and language-specific attributes.
Need help interpreting the samples?
Tell us which languages, services and workflows you want to evaluate. We will help map the sample files and technical specifications to your AI, search or document pipeline.
MADRID, SPAIN
Camino de las Huertas, 20, 28223 Pozuelo
Madrid, Spain
SAN FRANCISCO, USA
541 Jefferson Ave Ste 100, Redwood City
CA 94063, USA