Dataset metadata
Turkish Context-Sensitive Lemmatization Corpus
T-BDLD · Discover the scope, method, live derived counters, citation, and restricted-access policy of the T-BDLD Turkish corpus. Raw corpus files are not publicly distributed.
Scope and purpose
T-BDLD is a continuously expanding Turkish research corpus used to connect lexical entries with authentic contexts, inflected forms, collocations, valency, and domain or genre distributions.
Source classes
The processing pipeline separates books and uploaded documents, academic articles, theses, news, columns, and other reviewed web sources. These classes describe ingestion provenance; they do not grant redistribution rights.
Processing method
- Source identity and duplicate checks.
- Text extraction, OCR review, and noise filtering where required.
- Sentence segmentation, token alignment, and context-sensitive lemmatization.
- Incremental indexing and derived measurements in ClickHouse.
Temporal coverage
A reliable aggregate source-publication range is not yet available for every record, so no artificial start or end year is published. Corpus-ingestion timestamps are kept separate from source publication dates.
Query access and citation
The public interface returns bounded query results and derived measurements; it does not expose a bulk corpus export.
Fields deliberately omitted
No public distribution, download URL, DOI, ROR, licence, fixed dataset version, or unsupported temporal range is declared. Verified author ORCIDs identify people only; they do not change the corpus access policy. Omitted fields can be added only after independent verification and an explicit access decision.