Imported from sudoanmol/x-bookmarks-rag (
AGENTS.md). Install upstream withnpx skills add sudoanmol/x-bookmarks-rag. Copyright stays with the author.
x-bookmarks-rag
Natural-language search over the owner's X (Twitter) bookmarks. It captures bookmarks with a real browser, extracts the full text behind articles and links, embeds everything locally, and searches with hybrid retrieval.
No X API. Everything runs on this machine, except two remote services: Groq on its free tier, and Modal (owner's free credits) for GPU captions and transcripts.
This file is the handoff. Read all of it before you change anything. Many rules below cost hours to learn. The reasons are given so you can tell when a rule stops applying.
1. Rules you must not break
These come from the owner. They override your defaults.
| Rule | Why |
|---|---|
Use uv for everything: uv run, uv add, uv venv. |
Project standard. |
Add dependencies with uv add. Never hand-edit pyproject.toml. |
Keeps the lock file correct. |
Use trash, never rm. |
A hook blocks rm. |
Never pass -c user.email or -c user.name to git. |
The global git config is already correct. The owner had to clean up commits once because of this. |
Never commit secrets. .env holds GROQ_API_KEY. |
|
| Push only when the owner asks. The repo is local and has never been pushed. | |
| Do not add backward compatibility, fallbacks, or migrations. Remove the old path. | Owner's rule. See §7 for how schema changes are done here. |
| Do not build speculative abstractions. Build the simplest thing that fully meets the requirement. | Owner's rule: "refuse to solve problems we don't have". |
| Measure before you design. Look at the real data first. | Every good decision in this repo came from doing this. Every mistake came from skipping it. |
The owner writes and expects ASD-STE100 Simplified Technical English: short sentences, active voice, one meaning per word.
The session file
~/.config/x-bookmarks/state.json (mode 0600) holds a live X login. It sits
outside the repo on purpose.
- Never copy it into the repo.
- Never print its contents.
- Only X Article pages may use it. Every external page gets a browser context with no storage state, so a third-party site never sees the login.
2. Commands
uv run xbm login # opens a browser; the owner signs in by hand
uv run xbm sync # every stage below, in order, on what is new
uv run xbm status # coverage per stage, and readiness of each service
uv run xbm capture # capture new bookmarks since the watermark
uv run xbm inspect # report what the captured data contains
uv run xbm normalize # rebuild the DB from raw pages on disk
uv run xbm extract # fetch full text behind articles and links
uv run xbm caption # vision model on a Modal GPU (--limit, --retry)
uv run xbm transcribe # Whisper on Modal GPUs (--limit, --retry)
uv run xbm translate # Groq gpt-oss-120b over foreign chunks
uv run xbm index # chunk and embed (--rebuild, --no-* exclusions)
uv run xbm search "..." # hybrid search (-n, --author, --source)
uv run pytest -q # 137 tests, all offline, ~1s
Ollama must be running, with embeddinggemma pulled. That is the only model
the project needs.
3. Architecture
Data flows one way. Each stage only reads what the stage before it wrote.
X GraphQL -> capture -> raw pages (JSON on disk)
|
parse + normalize
|
bookmarks / authors / media / links
|
extract (Playwright + Defuddle) -> documents
caption (Modal + vLLM) -> captions
transcribe (Modal + Whisper) -> documents (kind video)
|
chunk -> embed -> chunks + chunk_vec + chunk_fts
|
search (RRF over vector + BM25)
| Module | Job |
|---|---|
config.py |
Paths, URLs, pacing. Loads .env by explicit path (see §8). |
session.py |
Saves and loads the browser storage state. |
capture.py |
Drives the bookmarks page and intercepts GraphQL responses. |
parse.py |
Turns a GraphQL page into rows. Network-free, so it is fully testable. |
normalize.py |
Rebuilds the DB from raw pages. |
db.py |
Schema and idempotent upserts. Owns all tables. |
extract.py |
Renders pages and pulls readable text with Defuddle. |
caption.py |
Modal app (vLLM on one H100) plus the local job and save logic. |
transcribe.py |
Modal app (faster-whisper on up to 8 L4s). Writes documents rows. |
translate.py |
Picks foreign chunks, translates them with Groq, caches by text hash. |
chunk.py |
Turns rows, documents, and captions into embeddable pieces. |
embed.py |
Ollama calls. Asymmetric prefixes (see §8). |
index.py |
Builds chunks, chunk_vec, chunk_fts. Owns the vec0 connection. |
search.py |
Hybrid retrieval with reciprocal rank fusion. |
mcp_server.py |
Stdio MCP tools for search, bookmark details, and document reading. |
inspect.py |
The gate report that drives build decisions. |
cli.py |
Typer entry point. |
Two connection functions. Do not mix them.
db.connect()— plain SQLite. Use for everything normal.index.connect()— loads the sqlite-vec extension. Any query touchingchunk_vecfails withno such module: vec0without it.
4. Data model
All tables live in db.py. data/bookmarks.db, WAL mode, busy_timeout=30000.
| Table | Notes |
|---|---|
authors |
author_id PK, screen_name, name, avatar_url, verified, description. Quoted-post authors are deliberately not stored. |
bookmarks |
tweet_id PK, sort_index, text, is_long, lang, quoted_text, article_id, article_title, article_preview, removed_at. |
media |
media_key PK, kind (photo/video/animated_gif), url, alt_text, duration_ms, bitrate, small_url/small_bitrate. |
links |
(tweet_id, url) PK, domain, title, description, from_card. |
documents |
url PK, kind (x_article/link/video), tweet_id (NULL for links), title, body (HTML), word_count, attempts, error. |
bookmark_documents |
View: which documents each bookmark has. A linked page belongs to every bookmark that links it, through links. Read documents through this, never documents.tweet_id. |
captions |
media_key PK, text, model, attempts, error. No FK to media: replace_media deletes and reinserts on every normalize. |
translations |
hash PK (of the original chunk text), source_text, text (NULL = already English), model. |
chunks |
id, tweet_id, source (post/quote/article/link/image/video), ref, position, text, source_text, lang, hash. |
chunk_vec |
vec0 virtual table, FLOAT[768]. |
chunk_fts |
fts5 external-content table over chunks. |
sync_state |
watermark = highest sort_index seen. |
raw_pages |
Every captured GraphQL page. The DB can always be rebuilt from these. |
Two design points worth keeping:
- Soft delete. Un-bookmarking sets
removed_at, so captions survive. sort_index, not post time. It orders by bookmark time, which is what incremental sync needs.
5. Current state (all verified, not estimated)
| Thing | Count |
|---|---|
| Counted on 2026-10-03. |
| Thing | Count |
|---|---|
| Bookmarks | 1,350 |
| Authors | 856 |
| Media | 949 — 606 photo, 335 video, 8 gif |
| Media with alt text | 18 |
| Links (unique) | 626 |
| Documents | 1,119 stored, 819 usable: 683 pages, 136 transcripts. Every link has a row. |
| Extracted words | 1,217,416 |
| Chunks / vectors | 5,005 / 5,005 (608 image, 708 video, 13 translated) |
| Captions | 605 of 606 photos. The one failure is a 404 at X. |
| Video | 335 clips, 51.2 hours. 136 have speech (533k words), 61 have audio but no speech, 133 have no audio track. |
Layers 1 and 2, image captions, and video transcripts are done and working. Search returns real passages from extracted pages, OCR text from images, and timestamped transcript text, not just post text.
The unusable documents are mostly "thin extraction" (under the word floor),
plus 10 binary targets (PDFs). links.url is stored under normalize_url(),
the same key as documents.url, so the two join directly.
6. Decisions already made. Do not re-open these.
Embedding model: embeddinggemma. Benchmarked against 4 alternatives on 55
generated queries over the real corpus.
| Model | Dims | R@1 | R@5 | MRR |
|---|---|---|---|---|
| embeddinggemma | 768 | 0.564 | 0.727 | 0.642 |
| qwen3-embedding:0.6b | 1024 | 0.509 | 0.818 | 0.626 |
| bge-m3 | 1024 | 0.455 | 0.727 | 0.583 |
| nomic-embed-text | 768 | 0.436 | 0.618 | 0.536 |
At 55 queries the 95% band is about ±0.13, so no model separates. The decision was made on speed: 3,437 chunks embed in about 2 minutes, where a 4B model needs 30 to 60. The benchmark measured dense retrieval alone; the real system also runs BM25, which narrows the gap further. The rejected models were deleted to free 17.4 GB.
Depth-1 crawling: none. The owner chose this after seeing the numbers:
5,342 outbound URLs, about 2.2 hours, and roughly 7x index growth, dominated by
GitHub repository navigation and documentation sidebars. Nothing is lost —
the outbound links are still inside documents.body, so this can be turned on
later without re-fetching.
No LLM answer layer. The owner will use this mostly through MCP, where the agent is already the model. Pre-summarizing would compress the evidence, and the agent would then summarize the summary. Retrieval returns passages; the caller reasons.
One vector space for all content types, rather than separate indexes.
Groq for enrichment, not embeddings. Groq serves no embedding models.
translate.py uses openai/gpt-oss-120b with reasoning_effort: low. The
free tier allows 8,000 tokens per minute, prompt and reply together, so text
goes in paragraph-aligned pieces of 1,500 characters and a 429 waits on
retry-after.
Translation is per chunk, cached by text hash. xbm translate builds the
same chunks xbm index would, and asks about a chunk when its X tag is not
English or its letters are mostly non-Latin. X's tag is wrong on most short
posts ("ro", "de", "pt" on English text), so the model replies ENGLISH for
those and the cache stores NULL. index.build puts the English in
chunks.text and the original in chunks.source_text. Of 35 candidates, 13
were really foreign: 4 Chinese and 3 Japanese posts or quotes, their
screenshots, a Chinese gist and repo page, a Japanese video, a Spanish post.
Captions run on Modal, not Groq. The Groq free tier (8,000 tokens per
minute, serial) needed days for 600 photos. Modal runs the same model,
Qwen/Qwen3.8-27B-FP8, under vLLM on one H100: 583 photos in one run, about
15 minutes of GPU and roughly $1. The app is ephemeral (app.run()), so it
stops when xbm caption exits. The owner pays for every GPU second: after
any Modal run, confirm modal container list --json prints [].
Transcripts run on Modal too, and are documents rows. faster-whisper
large-v3-turbo on up to 8 L4s did all 51.2 hours in under 6 minutes. This
replaced the planned Groq + local mlx-whisper split: an M3 Pro would need
about 2 hours. A transcript is a documents row with kind = 'video', keyed
by https://x.com/<handle>/status/<id>/video/<n>, with one <p> per minute
led by [m:ss]. So chunking, search, and read_document (paging and query)
needed no new code. READABLE_SQL exempts transcripts from the 60-word page
floor. A clip with no speech is stored with word_count = 0: done, never
served.
7. Hard-won lessons
Each of these cost real time. Do not rediscover them.
Playwright and X
- The GraphQL query ID changes on every X frontend deploy. Match only the
trailing operation name (
Bookmarks), never the full path. - Never make blocking calls inside a Playwright event handler in the sync
API. Collect
Responseobjects in the handler; read bodies in the main flow. - X throttles one session across concurrent browsers. The first parallel
run returned 61 of 153 articles empty. Every one succeeded on a serial
retry. Articles now use
ARTICLE_WORKERS = 1; links keepWORKERS = 4, because they spread over hundreds of hosts. - Never scroll an X Article page. The view mounts and unmounts surrounding
elements as you move, so scrolling makes extraction worse — measured at
542 words after two scrolls and 449 after four, against 478 with no scroll.
The GraphQL
content_stateis captured in parallel as the authoritative copy and proves the render was complete. bypass_csp=Trueis required. GitHub otherwise blocks script injection.- Playwright's sync API is bound to its creating thread. Each worker
thread must build its own
sync_playwright(), browser, and DB connection. - Rewrite arXiv
/pdf/to/abs/. A PDF URL makes the browser start a download instead of rendering.looks_like_a_file()guards the rest. - JavaScript apps can be empty at
domcontentloaded. Retry once onnetworkidlebefore calling an extraction thin.
Everything else
load_dotenv()must take an explicit path.find_dotenv()walks up from the calling frame, which breaks for scripts run from stdin.- Defuddle returns HTML, not markdown, even with
markdown: true. That is fine and deliberate:documents.bodykeeps HTML because the<a href>and<img src>inside are what later layers need. Markup comes off at chunk time inchunk.to_text(). - Schema changes: drop and rebuild, never migrate.
CREATE TABLE IF NOT EXISTSwill not add a column to an existing table, and the owner forbids migrations.raw_pagesmakes rebuilding safe. - SQLite:
LIMITgoes afterUNION ALL. Wrap each side in a subquery. - embeddinggemma uses asymmetric prefixes. Documents get
title: none | text:, queries gettask: search result | query:. Using the wrong one is a real handicap. Seeembed.py. - Verify test expectations against real output before trusting them. Two early test failures were wrong expectations, not wrong code.
- Cards name their target by t.co, so resolve the short link. Guessing which link a card belongs to attached cards to the wrong URL.
- Store the smallest MP4, not the best. Transcription wants speech, not pixels. A 44.5-minute clip is 3.3 GB at top bitrate and 81 MB at 256 kbps. Across 40 bookmarks this was 10.73 GB versus 0.27 GB, a 40x saving.
- Do not FK
captionstomedia.replace_mediadeletes and reinserts on every normalize.ON DELETE CASCADEwould wipe every caption onxbm sync.documentshas no FK for the same reason. - vLLM FP8 needs a CUDA devel image. DeepGEMM compiles kernels at
startup and asserts on a missing CUDA toolkit.
debian_slimhas none. Usenvidia/cuda:<ver>-devel, matched to the vLLM wheel (0.30 → CUDA 13). - Qwen3.8 is a hybrid Mamba model. vLLM's default
max_num_seqs=1024exceeds its Mamba cache blocks on an H100 (796). Set it to the batch size. - A raise in Modal
@entercrash-loops on the GPU while the client waits.caption.Model.loadkeeps the error and fails the first call instead, so the run ends and the app stops. - Finish the Modal
.map()generator. Leaving it unfinished closes it insideapp.run()and raisesaclose(): asynchronous generator is already running. Inzip, put the generator first. - Incremental sync stops one page after the watermark hit. An earlier version only stopped when a scroll returned no pages, so it walked all 69.
- faster-whisper 1.2.1 needs PyAV below 19. PyAV 19 removed the
metadata_errorsargument it passes toav.open. - faster-whisper's
decode_audiodrops invalid data without a word. A damaged download decodes to nothing and looks silent: a 2-hour podcast came back empty once. The container checks the byte count againstContent-Length, and treats audio with no detected language as an error. - A linked page has many owners. 28 URLs are bookmarked more than once.
Link documents used to carry the first bookmark's
tweet_id, so the others could not find or read their own page. Ownership now comes fromlinksthrough thebookmark_documentsview, and a shared page is chunked once per bookmark. - A capture with no media or links is the truth.
normalizereplaces them unconditionally, so a stale photo and its caption leave the index.
7b. The CLI
xbm syncruns capture, extract, caption, transcribe, translate, index. A failed stage does not stop the rest; the run exits non-zero at the end. Each stage is a_stage()function incli.pythat its own command andsyncboth call.- Exclusions persist in
~/.config/x-bookmarks/config.toml(exclude = [...],translate = true), read byconfig.settings().--no-*flags add to it for one run. An excluded source is skipped at its stage and dropped from the index (vectors too), but stays in the database. - GPU stages never retry failures on their own; only
--retrydoes. The usual failure is media deleted at X, and each try starts a GPU. synccaptions only at 20 new photos (caption.SYNC_MIN_PHOTOS). vLLM startup costs the same 5 to 8 H100 minutes for 1 photo as for 600.xbm captionignores the threshold. Transcripts need no threshold: L4 startup is about a minute.
8. MCP server
This is done. mcp_server.py uses MCP 2.0 and stdio transport. Codex has the
server registered as x-bookmarks.
Why three tools, not one
Search returns one passage per bookmark, so a long thread cannot crowd out a sharp short post. That hides scale:
| Chunks per bookmark | Bookmarks |
|---|---|
| 1 | 536 |
| 2–3 | 434 |
| 4–10 | 210 |
| 11+ | 31 |
The worst case is one bookmark whose linked Claude Code CHANGELOG is 78,729 words in 155 chunks. A caller sees 1 of 155. It needs a way to go deeper, and a way to know that deeper exists.
The tools
-
search_bookmarks(query, limit=10, author=None, source=None, after=None, before=None)Wrapssearch.search(). Return per hit:tweet_id,url,author,created_at,score,best_source,best_chunk(the passage that matched),best_ref(the page URL it came from),chunk_count,word_count,media,lang.chunk_countandword_countare what make tool 3 discoverable. Alsobest_source_text(the original of a translated passage) and, on the response,last_sync.afteris inclusive,beforeexclusive, on the post date.Filters apply before ranking (
search.allowed_chunks). They used to apply to the top 80 candidates, so--source video -n 3returned 2 hits. sqlite-vec'schunk_id IN (...)is a true pre-filter, except with one id: SQLite rewrites a one-itemINto=, and vec0 returns no rows for that inside a KNN query._vector_ranksshort-circuits the single-candidate case. -
get_bookmark(tweet_id)Full post text, quoted post, author, date, every link, every media item. Each media item includescaptionwhen one exists. Cheap and bounded. Do not include document bodies here. -
read_document(tweet_id_or_url, query=None, offset=0)The extracted page text. This one is dangerous if done naively: returning the CHANGELOG whole is roughly 100,000 tokens and would destroy the caller's context. So:- with
query: run the same hybrid search restricted to that document and return the best passages; - without: page through with
offsetand returnhas_more.
- with
Use stdio transport. Add a xbm-mcp entry point in [project.scripts].
Registration
Codex uses this global registration:
codex mcp add x-bookmarks -- uv run --project ~/Developer/x-bookmarks-rag xbm-mcp
The server was tested through an MCP stdio client. All three tools returned structured results against the live corpus.
9. Backlog after MCP
In the owner's chosen order. Each layer must leave a working product.
- Image captions and OCR. Done.
uv run xbm captionafter each sync captions only new photos, thenuv run xbm index. - Video transcripts. Done.
uv run xbm transcribeafter each sync, thenuv run xbm index. Usesmedia.small_url, neverurl. - Foreign language translation. Done.
uv run xbm translateafter each sync, thenuv run xbm index. - Web UI. Over the same
search.search()function.
10. How to verify your work
uv run pytest -q— 137 tests, all offline, about one second. Keep it that way. Tests use synthetic GraphQL fixtures intests/fixtures.pyand stubembed.embed_documents/embed.embed_querywith a deterministic vector.- Run a real query and read the passages:
uv run xbm search "what was the article about kv caching". A good result shows the matching passage and afrom <url>line, not the post text. - Never report a step as done without running it. The owner checks.
