Imported from sadit/TextSearch.jl (
AGENTS.md). Install upstream withnpx skills add sadit/TextSearch.jl. Copyright stays with the author.
AGENTIC.md
Guidance for AI coding agents (Claude Code and similar) working in this repository.
What this package is
TextSearch.jl is a Julia package that turns text into vector/sparse representations
(BOW, TF, TF-IDF, entropy-based weightings, BM25) and provides an inverted-file index
(via InvertedFiles.jl) for searching them.
It is meant to be paired with SimilaritySearch.jl.
The package was previously named TextModel.jl.
Processing pipeline (mental model)
Understanding the data flow makes it much easier to find where a change belongs:
raw text/corpus
→ TextConfig (tokenizer/textconfig.jl) preprocessing options (lowercase, diacritics, url/user/number
grouping, q-grams, n-grams, skip-grams, token transforms)
→ normalize_text (tokenizer/normalize.jl) character-level normalization
→ tokenize (tokenizer/tokenize.jl) produces a TokenizedText (list of token strings)
→ Vocabulary (voc.jl, updatevoc.jl) token ⇄ id mapping, occurrence/doc-frequency counters
→ BOW / bagofwords (bow.jl) Dict{UInt32,Int32} bag-of-words per document
→ VectorModel (vmodel.jl, emodel.jl) local/global weighting schemes → SparseVector{Float32,Int32}
or
→ BM25 (bm25/BM25.jl) + BM25InvertedFile (bm25/invfile.jl, bm25/invfilesearch.jl) index + kNN search
Tokenization (textconfig.jl, tokentrans.jl, normalize.jl, tokenize.jl) lives under
tokenizer/ as the internalized TextSearch.Tokenizer submodule; BM25 (bm25.jl,
bm25invfile.jl, bm25invfilesearch.jl) lives under bm25/ as TextSearch.BM25 (the scorer
struct is BM25Scorer, since a module can't share its name with a binding inside it);
InvertedFiles/Intersections are internalized the same way (SimilaritySearch.InvertedFiles/
SimilaritySearch.Intersections, no longer external deps).
sparseconversions.jl / dvec.jl bridge between the package's Dict-based BOW
(Dict{UInt32,Int32}) and SparseArrays/SimilaritySearch vector types (arithmetic, norms,
distances). The final weighted vectors produced by vectorize/vectorize! are
SparseVector{Float32,Int32}, not a Dict — the old SVEC = Dict{UInt32,Float32} alias was
removed as dead code once nothing in src/ still produced or consumed it.
Source file map
| File | Responsibility |
|---|---|
TextSearch.jl |
Module entry point, includes, BOW type alias |
tokenizer/textconfig.jl |
TextConfig, Skipgram — tokenization/preprocessing configuration |
tokenizer/tokentrans.jl |
AbstractTokenTransformation hooks (stemming, stopwords, chaining) |
tokenizer/normalize.jl |
Character-level text normalization, emoji detection |
tokenizer/tokenize.jl |
TokenizedText, tokenize/tokenize_corpus, q-grams/n-grams/skip-grams |
tokenizer/generators.jl |
AbstractTokenGenerator and built-in generators (unigram, word n-gram, character q-gram) |
voc.jl |
Vocabulary type: token↔id table, occurrence/ndocs counters |
updatevoc.jl |
Merging/updating Vocabulary instances |
approxvoc.jl |
Approximate vocabulary lookup (QgramsLookup) for fuzzy/OOV matching |
bow.jl |
Bag-of-words construction from tokenized text |
sparseconversions.jl |
Conversions between Dict sparse vectors and SparseArrays |
dvec.jl |
Arithmetic/distance operations (+, -, dot, norm, centroid, Cosine/Angle) on SparseVector (and generic Dict{Ti,Tv} arithmetic) |
vmodel.jl |
VectorModel, local/global weighting schemes (TF, IDF, TP, binary), vectorize |
emodel.jl |
Entropy-based weighting schemes (EntropyWeighting, CombineWeighting) |
multi.jl |
Merging/joining VectorModels (update!, joinmodel) — requires KCenters |
bm25/BM25.jl |
BM25Scorer scoring struct and bm25score/tokenscore |
bm25/invfile.jl |
BM25InvertedFile — the inverted-file index built on InvertedFiles.jl |
bm25/invfilesearch.jl |
kNN search over BM25InvertedFile (uses Intersections.jl) |
deprecated.jl |
Backwards-compatible shims for renamed/removed APIs |
Dev environment
# from repo root
julia -t 3 --project=. -e 'using Pkg; Pkg.instantiate()'
Run the test suite:
julia -t 3 --project=. -e 'using Pkg; Pkg.test()'
# or, from a REPL with the project active:
] test
Individual test files live under test/ and are included from test/runtests.jl:
tok.jl (tokenization), voc.jl (vocabulary), vec.jl (vector models/weighting),
search.jl (BM25 inverted file search).
Build the docs (Documenter.jl):
julia --project=docs docs/make.jl
CI (.github/workflows/ci.yml) runs on Julia 1.12 for Ubuntu and Windows and uploads
coverage to Codecov. documentation.yml builds and deploys docs on pushes/tags to main.
Conventions and gotchas
- Julia version floor is 1.10 (
Project.toml[compat] julia = "^1.10"). Don't use syntax/stdlib features newer than that. Aqua.jlruns in the test suite (test/runtests.jl): ambiguity and type-piracy checks (Aqua.test_ambiguities,Aqua.test_piracies). If you add methods that extendBase/LinearAlgebra/SparseArraysfunctions on non-owned types, add them to thetreat_as_ownlist inruntests.jlif they're intentional, otherwise avoid the piracy.BOWis a plainDict(Dict{UInt32,Int32}, defined inTextSearch.jl) — a short-lived, incrementally-built intermediate. The final weighted vectors (vectorize/vectorize!) areSparseVector{Float32,Int32}, not aDict— arithmetic/ distance operators for both live indvec.jl— extend there, not ad hoc elsewhere.- Buffer pooling: separate channel-based per-thread pools avoid allocations —
BOW_CACHES(aChannel{BOW}) invoc.jlbacksVocabularybuilding only;bagofwords/bagofwords!/bagofwords_corpus(bow.jl) take/create a plainBOWdirectly (no wrapper struct),sizehint!ed up front fromvoc'savgdoclenvia_bow_sizehint;VectorizeBuffer/VECTORIZE_CACHESinvmodel.jlbacks the performance-sensitivevectorize!path (sorts/RLEs token ids directly, noDictinvolved); andTokenizerBuffer/TOKENIZER_CACHESinside theTokenizermodule (usetokenizerbuffer(f)) backs tokenization scratch space, borrowed on demand by the others rather than duplicated. - Parallelism uses SimilaritySearch's
@BATCHES(v1.0+; noPolyesterdependency anywhere in this package anymore). Simple per-item loops use the one-argument form (@BATCHES getminbatch(n) for i in 1:n ... end);voc.jl'stokenize_and_append!instead counts each batch into its own lock-free table and folds the tables pairwise, in corpus order, so token ids come out in order of first appearance whatever the thread count. Always callgetminbatch(n)with one argument —getminbatch(n, nt)'s second argument is a thread count, not a second corpus-size-like quantity; passing anything corpus-derived there silently returns a useless batch size (this was a real, previously-unnoticed bug in this codebase before the v1.0 migration, since the loops using it wereThreads.@threads, which doesn't even take aminbatch).- Known residual risk:
InvertedFileContext's search-time scratch buffers (positions,cont_u32,cont_iw,cont_iiw,knns, sized byThreads.maxthreadid()) are still indexed byThreads.threadid()(getcontainer/getpositionsininvertedfiles/invfile.jl). This is unchanged by the v1.0 migration (SimilaritySearch's own internal state moved to@batchid()-indexing instead, precisely to avoid this class of bug — see itsAGENTS.md/set_batch_scheduler!docs), but TextSearch's version wasn't ported becauseInvertedFileContextis a single global cached singleton (DEFAULT_CACHE_INVFILES[], returned bygetcontext) shared across whatever parallelizes concurrentsearch/push_item!calls outside TextSearch's control — a clean@batchid()fix needs that caller to tag a per-batch context copy (@set ctx.batchid = @batchid()) before calling in, which isn't a contract TextSearch can enforce today. Only touch this if you're prepared to either add that per-call-site tagging contract, or redesign away from one-shared-context-with-N-thread-slots entirely (e.g. one context per concurrent caller, no internal arrays at all).
- Known residual risk:
- Struct immutability: most core types (
TextConfig,Vocabulary,BM25, …) are immutablestructs with explicit copy-constructors (e.g.TextConfig(c::TextConfig; kwargs...)) for "update a field" patterns — follow that pattern instead of adding mutability. - Custom
Base.showmethods exist for most structs with aprefix/indentkeyword convention for nested pretty-printing — match this style for new types. - Docstrings use standard Julia docstring format with a signature block followed by
a description and
- field: descriptionbullet lists; keep new public API documented the same way and cross-reference with[`Type`](@ref). - This package depends on unreleased/companion packages from the same author
(
InvertedFiles.jl,Intersections.jl,SimilaritySearch.jl) — check their APIs when a search/index-related change looks like it needs upstream changes too.
Where to look for more context
README.md— install/usage overview, links to Pluto notebook examples inexamples/.docs/src/index.md,docs/src/api.md— narrative docs and API reference source.examples/invindex.jl,examples/searchgraph.jl— Pluto notebooks demonstrating end-to-end usage.
