Imported from evrenesat/asky (
src/asky/research/AGENTS.md). Install upstream withnpx skills add evrenesat/asky --skill research. Copyright stays with the author.
Research Package (asky/research/)
RAG-powered research mode with caching, semantic search, and persistent memory.
Module Overview
| Module | Purpose |
|---|---|
tools.py |
Research tool schemas and executors |
cache.py |
ResearchCache for URL content/links |
vector_store.py |
VectorStore for hybrid semantic search |
vector_store_chunk_link_ops.py |
Chunk/link embedding and retrieval operations |
vector_store_finding_ops.py |
Research findings embedding/search operations |
vector_store_common.py |
Shared vector math and constants |
embeddings.py |
EmbeddingClient for local embeddings |
chunker.py |
Token-aware text chunking |
source_shortlist.py |
Pre-LLM source ranking pipeline |
query_classifier.py |
Query classification for one-shot summarization |
query_expansion.py |
Decomposing queries into sub-queries. |
evidence_extraction.py |
Post-retrieval LLM fact extraction. |
shortlist_collect.py |
Candidate/seed-link collection stage |
shortlist_score.py |
Semantic + heuristic scoring stage |
shortlist_types.py |
Shared shortlist datatypes and callback aliases |
corpus_context.py |
Corpus-aware context extraction for shortlist query enrichment |
sections.py |
Deterministic section indexing + strict matching |
adapters.py |
Local-source loading with plugin extension support |
Research Tools (tools.py)
Available Tools
| Tool | Description |
|---|---|
extract_links |
Cache URL content, return discovered links |
get_link_summaries |
Get AI-generated page summaries |
get_relevant_content |
RAG retrieval of relevant chunks |
get_full_content |
Complete cached content |
list_sections |
List section headings for local corpus |
summarize_section |
Deep summary for one local section |
save_finding |
Persist insights to research memory |
query_research_memory |
Semantic search over saved findings |
Tool schemas also support optional system_prompt_guideline metadata used by
chat/system-prompt assembly when the tool is enabled for a run.
When a chat session is active, registry plumbing can inject session_id into
memory tool calls so findings are written/read in session scope.
Research chat flow now guarantees that a session exists (auto-created when needed),
so session-scoped memory isolation is available by default in research mode.
get_relevant_content and get_full_content now also accept internal
corpus_urls identifiers and can resolve safe cache handles
(corpus://cache/<id>) to cached entries. This enables local-corpus retrieval
without exposing filesystem paths to the model.
Section-scoped retrieval contract:
- Preferred:
section_ref(corpus://cache/<id>#section=<section-id>) or explicitsection_id. - Compatibility: legacy
corpus://cache/<id>/<section-id>source suffixes are accepted. list_sectionsdefaults to canonical body sections and emitssection_reffor each row.- CLI boundary note: positional
--summarize-section <value>is a title/query strict-match input (SECTION_QUERY), not asection_id. - CLI wrappers that need deterministic section targeting must pass
section_idexplicitly (for example via--section-id).
Parameter priority in get_relevant_content:
- If both
section_refandsection_idare provided,section_reftakes precedence. - If only
section_idis provided alongside a source URL, it is applied as a section filter on that source.
Tool Sets by Stage
Tools are grouped into constants to support per-stage exposure:
ACQUISITION_TOOL_NAMES:extract_links,get_link_summaries,get_full_content.RETRIEVAL_TOOL_NAMES:get_relevant_content,list_sections,summarize_section,save_finding,query_research_memory.
When a corpus is pre-loaded by acquisition stages, acquisition tools are excluded to prevent redundant LLM work.
Execution Flow
extract_links(urls, query?)
↓
ResearchCache.cache_url()
↓
VectorStore.store_chunk_embeddings()
↓
get_relevant_content(urls, query)
↓
Hybrid ranking (dense + lexical)
↓
Diverse chunk selection
ResearchCache (cache.py)
Caches fetched URL content and extracted links with TTL.
Scope note:
research_cacheis global to the activeDB_PATH(not session-bound).- Session cleanup (
session clean-research) does not directly purgeresearch_cache,content_chunks, orlink_embeddings; those are removed by TTL expiry cleanup (or explicit future purge tooling).
Key Features
- TTL-based expiry: Configurable
cache_ttl_hours - Startup cleanup: Expired entries purged on init (daemon thread)
- Content + links: Both cached together per URL
- Invalidation: Clears related vectors when content changes
- On-demand summarization:
get_link_summariesperforms synchronous summarization if the cache entry is missing or stale. - Refresh on re-ingest: Re-caching the same source key refreshes expiry/content by design.
Schema (SQLite)
research_cache: URL, content, links JSON, timestamps, TTLresearch_findings: Persistent research insights with embeddings
ResearchCache also exposes helper lookups used by manual/debug retrieval
workflows:
get_cached_by_id(cache_id)list_cached_sources(limit)
VectorStore (vector_store.py)
Hybrid semantic search combining ChromaDB dense and SQLite BM25 lexical retrieval.
Collections
| Collection | Content |
|---|---|
content_chunks |
Text chunks from cached pages |
link_embeddings |
Link anchor text for relevance filtering |
research_findings |
Saved insights for memory queries |
Hybrid Ranking
final_score = (dense_weight * semantic_score) + ((1 - dense_weight) * lexical_score)
- Dense: Chroma nearest-neighbor cosine similarity
- Lexical: SQLite FTS5 BM25 scoring
- Fallback: SQLite-based cosine scan if Chroma unavailable
Key Methods
store_chunk_embeddings(): Generate and persist chunk vectorssearch_content_chunks(): Hybrid search with diversity filteringsearch_relevant_links(): Filter links by semantic relevanceclear_cache_embeddings(): Remove stale vectors on invalidationdelete_findings_by_session(): Session-scoped cleanup of findings and vectors
Internal Module Split
vector_store.pynow focuses onVectorStorelifecycle, DB/Chroma capability checks, and compatibility wrappers.- Heavy operations were extracted:
vector_store_chunk_link_ops.pyfor content chunks and linksvector_store_finding_ops.pyfor research memory findingsvector_store_common.pyfor shared math/constants
- Chroma chunk/link query filters use single-operator metadata conditions
(
$and) for compatibility with stricter Chroma metadata parsing.
EmbeddingClient (embeddings.py)
Local sentence-transformer embeddings.
Configuration
- Model:
all-MiniLM-L6-v2(default) - Device: CPU or CUDA
- Batch size: Configurable for memory management
Features
- Singleton pattern for efficient reuse
- Lazy loading with cache-first Hugging Face download
- Pre-encode truncation to model max sequence length for embedding inputs
- Token counting for chunk alignment
- Usage stats exposed for banner display
Text Chunker (chunker.py)
Token-aware sentence chunking for optimal embedding boundaries.
Strategy
- Split text into sentences
- Build chunks within token budget
- Maintain overlap for context continuity
- Char-based fallback for non-sentence text
Query Classifier (query_classifier.py)
Analyzes user queries to detect one-shot summarization requests and determine appropriate response mode.
Classification Modes
- one_shot: Direct summarization without clarification questions (small corpus + clear intent)
- research: Complex workflow requiring clarification (large corpus or vague query)
Decision Tree
- Force research mode override (config) → research
- Empty corpus → research
- Vague query → research
- Corpus > threshold → research
- Has summarization keywords AND corpus ≤ threshold → one_shot
- Otherwise → research (safe default)
Key Features
- Keyword detection: Primary (summarize, summary, overview) and secondary (key points, main ideas, tldr)
- Vague query detection: Short queries (<10 chars) and generic phrases
- Confidence scoring: 0.0-1.0 based on detection factors
- Configurable thresholds: Default 10 documents, aggressive mode 20 documents
- Deterministic: Same inputs always produce same classification
Configuration
[query_classification]
enabled = true
one_shot_document_threshold = 10
aggressive_mode = false
force_research_mode = false
Integration
Called during preload pipeline (api/preload.py) after local ingestion completes. Classification result is stored in PreloadResolution.query_classification and used by prompt construction to adjust system guidance.
Source Shortlist (source_shortlist.py)
Pre-LLM source ranking to improve prompt relevance.
Pipeline
- Extract URLs from prompt
- Optional keyphrase extraction (YAKE)
- Collect candidates: seed URLs + search results + expanded links
- Fetch and extract main content
- Score with embeddings + heuristics
- Select top-k diverse sources
shortlist_prompt_sources(...) now accepts an optional trace_callback and
forwards it across search/fetch/seed-link transport paths when supported. This
enables verbose metadata traces for shortlist-stage HTTP calls (including
request failures such as 401/403).
Seed URL extraction accepts both explicit http(s)://... links and bare-domain
targets (e.g., example.com/path), normalizing bare targets to https://...
for deterministic fetch behavior.
When formatting shortlist context, explicitly mentioned seed URLs are included in the model-facing shortlist context even if they rank below the usual top-k context cutoff.
Internal Module Split
source_shortlist.pykeeps public API and orchestration.shortlist_collect.pyhandles candidate gathering and seed-link expansion.shortlist_score.pyhandles embedding-based scoring and ranking reasons.shortlist_types.pyholds shared shortlist datatypes and callback aliases.
Enablement
- Global flags per mode (research/standard)
- Per-model override in
models.toml --leanflag disables for single run- Default shortlist budgets are bounded (
max_candidates=40,max_fetch_urls=20) and remain configurable inresearch.toml. - In API preload policy resolution for local-corpus turns:
research_source_mode=local_onlyis a hard disable path for shortlist.research_source_mode=mixeduses intent routing:- deterministic
webintent -> shortlist enabled - deterministic
localintent -> shortlist disabled - ambiguous intent -> optional interface-model decision, fail-safe disabled fallback.
- deterministic
Shortlist behavior matrix (effective runtime):
- Standard mode: policy-driven (lean/request/model/global).
- Research
web_only: policy-driven (lean/request/model/global). - Research
local_only: always disabled to keep corpus-only contract. - Research
mixed: adaptive intent policy over local+web candidates.
Operational note:
- Shortlist enablement governs only pre-LLM candidate preparation; it does not guarantee subsequent deep page fetch tool calls by the main model.
Local Loader (adapters.py)
adapters.py provides local-source loading and can be extended by plugins via the LOCAL_SOURCE_HANDLER_REGISTER hook.
- Accepted targets:
local://...,file://..., absolute/relative local paths. - Local loading is enabled only when
research.local_document_rootsis configured. - Absolute paths are accepted only when they are inside configured roots (unless
research.allow_absolute_paths_outside_rootsis true). - Ingested files can optionally be restricted to a configured global allowlist (
research.allowed_ingestion_extensions). - Root-relative targets (for example
/nested/doc.txt) resolve under configured roots. extract_local_source_targets(...)provides deterministic token extraction from prompts for pre-LLM local preload.- Directory targets (discover): produce file links as
local://...(non-recursive in v1). Directory itself is not ingested as a pseudo-document and does not count towards document totals. - File targets (read/discover): normalize to plain text for cache/indexing.
- Text-like:
.txt,.md,.markdown,.html,.htm,.json,.csv - Document-like:
.pdf,.epubvia PyMuPDF - Plugin-provided: any extension registered by plugins (e.g. audio/image via transcriber plugins).
- Text-like:
- Directory targets are discovery-only in v1; select returned file links for content reads.
Local Target Guardrails in Research Tools
- Generic research LLM tools (
extract_links,get_link_summaries,get_relevant_content,get_full_content) reject local filesystem targets. - Safe corpus handles (
corpus://cache/<id>) are allowed for retrieval tools and map to cached entries without revealing local paths. list_sectionsandsummarize_sectionare local-corpus-only tools:- web URLs are rejected explicitly,
- they are hidden from registry exposure in
web_onlymode, - in
mixedmode they only accept corpus handles (corpus://cache/<id>). summarize_sectionresolves aliases to canonical body sections and refuses tiny sections.
- This prevents implicit local-file access via broad URL-oriented tools.
- Local-file access should be handled through explicit local-source tooling/adapters in dedicated workflows.
- Query preprocessing can redact local path tokens from model-visible user text when local preload is active, avoiding direct file-path exposure to the model.
Dependencies
research/
├── tools.py → cache.py, vector_store.py
├── cache.py → embeddings.py (for summaries)
├── vector_store.py → embeddings.py, chunker.py
├── source_shortlist.py → retrieval.py, embeddings.py
└── adapters.py → builtin local loading helpers
