Imported from mkubicek/xTap (
AGENTS.md). Install upstream withnpx skills add mkubicek/xTap. Copyright stays with the author.
AGENTS.md - xTap
What This Is
xTap is a browser extension (Chrome + Firefox) that passively captures tweets from X/Twitter by intercepting GraphQL API responses the browser already receives. No scraping, no extra requests — just structured JSONL output of what the user sees.
Repo: github.com/mkubicek/xTap License: MIT (public repo)
Architecture
content-main.js (MAIN world)
│ Patches fetch() + XHR.open() to intercept GraphQL responses
│ Emits CustomEvent with random per-page name
▼
content-bridge.js (ISOLATED world)
│ Reads event name from <meta> tag, listens, relays
│ Removes <meta> immediately after reading
▼
background.js (Service Worker, ES module)
│ Parses tweet data via lib/tweet-parser.js
│ Deduplicates (Set of seen IDs, max 50k; session storage in dev, local in prod)
│ Batches (50 tweets or 30–45s jittered flush)
│ Debug logging: intercepts console.log/warn/error, sends to host
│ Transport: HTTP daemon only (reprobes on failure with 30s cooldown)
▼
┌─── HTTP transport ─────────────────────────────────────────────┐
│ xtap_daemon.py (127.0.0.1:17381) │
│ Managed by launchd (macOS), systemd (Linux), Scheduled Task │
│ (Windows). Bearer token auth from ~/.xtap/secret │
│ Endpoints: GET /status, POST /tweets, /log, /test-path, │
│ /check-ytdlp, /download-video, /download-status │
│ /tweets accepts image_download:true to opt-in to background │
│ image fetches; returns images_queued in the response. │
└────────────────────────────────────────────────────────────────┘
┌─── Native messaging (token bootstrap only) ───────────────────┐
│ xtap_host.py (Python, stdio) │
│ Browser native messaging protocol (Chrome/Firefox) │
│ Serves GET_TOKEN to bootstrap HTTP transport │
│ Crashes logged to ~/.xtap/host-error.log │
└────────────────────────────────────────────────────────────────┘
│ xtap_daemon.py uses shared logic from xtap_core.py
▼
tweets-YYYY-MM-DD.jsonl (daily rotation)
debug-YYYY-MM-DD.log (when debug logging enabled)
media/<tweet_id>/* (when image download enabled)
media-manifest.jsonl (when image download enabled — append-only download log)
Key Design Decisions
- Two content scripts (MAIN + ISOLATED): MV3 requires this split. MAIN world can patch browser APIs but can't use chrome.runtime. ISOLATED world bridges the gap.
- Random event channel: The CustomEvent name is generated per page load (
'_' + Math.random().toString(36).slice(2)) and passed via a<meta>tag that's immediately removed. Avoids predictable DOM markers. - HTTP-only transport: All data flows through the HTTP daemon (
xtap_daemon.py), managed by launchd (macOS), systemd (Linux), or Scheduled Task (Windows). On macOS, it runs outside browser TCC sandboxes, allowing writes to protected paths. If the daemon goes down, tweets are buffered in memory and the extension reprobes every 30 seconds until the daemon recovers. - Token bootstrap via native messaging: On first run, the extension connects to
xtap_host.pyvia native messaging to requestGET_TOKEN(reads~/.xtap/secret). The token is cached inchrome.storage.localfor subsequent HTTP requests. The native host handles nothing else — all data goes through HTTP. - Shared core logic:
xtap_core.pycontains all file I/O logic (load seen IDs, resolve output dir, write tweets/logs, test path), used byxtap_daemon.py. - Environment detection:
isDevMode = !chrome.runtime.getManifest().update_url— packed CWS extensions haveupdate_url, unpacked don't. Used to switch seenIds storage between session (dev) and local (production). - Volatile dev cache: In dev mode (unpacked),
seenIdsis stored inchrome.storage.session, which clears on extension reload. This eliminates the need to manually clear storage during development. Production behavior is unchanged (persisted tochrome.storage.local). - Dedup in service worker: Multiple tabs feed the same service worker.
seenIdsSet (max 50,000, FIFO eviction) prevents duplicates. In production, persisted tochrome.storage.localacross sessions; in dev mode, uses volatilechrome.storage.session. Both host and daemon also load seen IDs from existing JSONL files on startup. - Image backfill bypass: When image download is enabled, duplicate tweets with photo URLs may be forwarded to the daemon once per service-worker session so skipped images can be downloaded. These forwards emit
IMAGE_BACKFILLtrace events and do not incrementsessionCountorallTimeCount; toggling image download from off to on clears the session-only backfill cache so tweets can be rechecked. - Jittered flush: Batch flush uses
setTimeoutwith randomized interval (30s base + up to 50% jitter = 30–45s), re-randomized each cycle. Avoids clockwork-regular patterns. - Path validation: When the user sets a custom output directory, the service worker sends a
TEST_PATHrequest to the HTTP daemon, which attemptsmakedirs+ write/delete of a temp file before accepting the path. - Error resilience: The native host logs crashes to
~/.xtap/host-error.logwith Python version and traceback. The HTTP daemon returns error status codes and logs startup diagnostics (Python version, output dir, token status). When the daemon is unreachable, the extension shows a red "!" badge and buffers tweets until the next successful reprobe (30s cooldown). The popup auto-refreshes every 2 seconds to reflect transport state changes. - Daemon debug logging: Set
XTAP_LOG_LEVEL=debugto get per-request logging (method, path, duration, tweet counts, tracebacks). Configured via environment variable in the service template (launchd/systemd). Re-runinstall.shafter changing. - Image download tuning (env vars, all optional):
XTAP_IMAGE_DELAY_MS(default 100) sets the inter-request delay for the background image worker.XTAP_MAX_FILE_MB(default 50) caps the size of a single image; the worker aborts mid-stream and deletes the.partfile when exceeded.XTAP_MAX_MEDIA_MB(default unset = unlimited) caps the cumulative bytes downloaded per process; further jobs logskipped:quota. Bad values fall back to defaults with a stderr warning instead of crashing the daemon. - Video download tuning (env vars, optional):
XTAP_MAX_VIDEO_MB(default 500) caps the direct MP4 fallback download size and deletes the.partfile if exceeded; set0or less to disable the cap. Bad values fall back to the default with a stderr warning. - Daemon connection lifecycle (env vars, optional):
XTAP_CONN_TIMEOUT_S(default 15) is the per-connection socket timeout onDaemonHandler— required because handler threads are non-daemon (daemon_threads = False) andserver_close()joins them at shutdown; without it one idle connection blocks shutdown until SIGKILL. The timeout is per blocking read, not a total request deadline, soXTAP_SHUTDOWN_GRACE_S(default 10) additionally bounds the shutdown join and force-exits if a trickle peer keeps a handler alive. Both always resolve to a finite positive value — bad or<=0input falls back to the default.
Stealth Constraints
These are non-negotiable. xTap must remain completely passive.
- Zero extra network requests — never fetch, POST, or call any X/Twitter endpoint. The extension only reads responses the browser already received.
- Native-looking patches —
toString()on patchedfetchreturns'function fetch() { [native code] }'.XHR.opentoString returns the original native string.fetch.nameis set to'fetch'viaObject.defineProperty. - No expando properties — XHR URL tracking uses a
WeakMap, never attaches properties to instances. - No DOM footprint — no injected elements, no visible page modifications. The only transient artifact is the
<meta name="__cfg">tag, removed within milliseconds by the bridge script. - No console output in page context — all logging happens in the service worker, which runs outside the page's JavaScript environment.
- Minimal permissions — only
storage,nativeMessaging, andalarms. Host permissions scoped tox.com,twitter.com, and127.0.0.1(local daemon only). NowebRequest, notabs, noscripting, no web-accessible resources. The debug dashboard is an internal extension page (chrome-extension://origin), not a web-accessible resource. - Random event channel — per-page-load name, meta tag removed immediately after reading.
- Only
open()patched on XHR —send()is not patched, so non-GraphQL XHR calls have clean stack traces.
Any change that adds network requests to X/Twitter domains must be rejected.
Carve-outs (opt-in only, daemon-side, off by default):
- Video download (
/download-video) — user-initiated via popup button; requires anhttps://x.com/.../status/<id>orhttps://twitter.com/.../status/<id>URL; callspbs.twimg.com/video.twimg.comfor direct MP4 fallback and runs yt-dlp when available. - Image download (
image_download:trueon/tweets) — opt-in toggle in popup. The daemon fetchespbs.twimg.comphotos in a single background worker. Hostname allowlist enforced; redirects blocked; per-file size cap; never enabled by default.
The browser-side capture path stays passive. These carve-outs run on the daemon, not the extension, so the page itself never originates the request.
File Structure
xTap/
├── manifest.json # Chrome MV3 manifest (permissions: storage, nativeMessaging)
├── manifest.firefox.json # Firefox MV3 manifest (generated — do not edit)
├── background.js # Service worker (ES module) - transport, parsing, dedup
├── content-main.js # MAIN world - fetch/XHR patching
├── content-bridge.js # ISOLATED world - event relay
├── popup.html/js/css # Extension popup (stats, pause/resume, output dir)
├── debug.html/js/css # Debug dashboard (live events, transport health, debug/discovery toggles, parser sandbox)
├── icons/ # Extension icons (16, 48, 128)
├── lib/
│ └── tweet-parser.js # GraphQL response → normalized tweet objects
└── native-host/
├── xtap_core.py # Shared file I/O logic (used by host + daemon)
├── xtap_host.py # Native messaging host — token bootstrap only (Python, stdio)
├── xtap_daemon.py # HTTP daemon (127.0.0.1:17381)
├── com.xtap.daemon.plist # launchd plist template (macOS)
├── com.xtap.daemon.service # systemd unit template (Linux)
├── com.xtap.host.json # Native messaging host manifest (Chrome)
├── com.xtap.host.firefox.json # Native messaging host manifest (Firefox)
├── install.sh # macOS/Linux installer (+ daemon)
├── install.ps1 # Windows installer (+ daemon)
├── xtap_host.bat # Windows native host wrapper
└── xtap_daemon.bat # Windows daemon wrapper
Supported Endpoints
The tweet parser (lib/tweet-parser.js) has known instruction paths for:
HomeTimeline, HomeLatestTimeline, UserTweets, UserTweetsAndReplies, UserMedia, UserLikes, TweetDetail, SearchTimeline, ListLatestTweetsTimeline, Bookmarks, Likes, CommunityTweetsTimeline, BookmarkFolderTimeline
TweetResultByRestId is also handled — it returns a single tweet (not a timeline). Full article payloads from this endpoint are captured when content_state.blocks is present; article stubs without the body are skipped so they do not poison dedup.
Unknown endpoints fall back to a recursive search for instructions[] arrays (max depth 5). Non-tweet endpoints are filtered in background.js via IGNORED_ENDPOINTS.
Output Schema
Each JSONL line contains:
{
"id": "1234567890",
"url": "https://x.com/handle/status/1234567890",
"created_at": "2024-01-01T00:00:00.000Z", // ISO 8601
"author": {
"id": "987654321",
"username": "handle",
"display_name": "Display Name",
"verified": false,
"is_blue_verified": true,
"follower_count": 1234
},
"text": "Full tweet text...",
"lang": "en",
"metrics": {
"likes": 10, "retweets": 5, "replies": 2,
"views": 1000, "bookmarks": 1, "quotes": 0
},
"media": [{"type": "photo|video|animated_gif", "url": "...", "alt_text": "...", "duration_ms": 1234}],
"urls": [{"display": "...", "expanded": "...", "shortened": "..."}],
"hashtags": ["tag"],
"mentions": [{"id": "...", "username": "..."}],
"in_reply_to": null,
"quoted_tweet_id": null,
"conversation_id": "1234567890",
"is_retweet": false,
"retweeted_tweet_id": null,
"is_subscriber_only": false,
"is_article": true, // only for long-form articles
"article": { // only for long-form articles
"title": "Article Title",
"text": "Rendered plain text with  refs",
"blocks": [], // raw Draft.js content_state blocks
"media": [{
"id": "...",
"url": "https://pbs.twimg.com/...",
"filename": "image.png",
"local_path": "media/<tweet_id>/image.png",
"width": 1200, "height": 800
}]
},
"source_endpoint": "HomeTimeline",
"captured_at": "2024-01-01T00:00:00.000Z"
}
Notes: media[].duration_ms only present for videos. views may be null. For retweets, text contains the full original tweet text (not the truncated RT @user: form). For articles, is_article and article are present — article.text is a markdown rendering with inline  image refs, article.blocks preserves the raw Draft.js structure, and article.media[] lists images with CDN URLs and local paths. Article stubs without content_state.blocks are skipped rather than saved as incomplete rows, so a later full article payload can still be captured. Top-level media[] entries deliberately do NOT carry local_path — when image download is enabled, photos land at media/<tweet_id>/<basename(url)> by convention. Consumers reconstruct the path; the daemon never round-trips it through the JSONL.
Known Issues
macOS TCC (Transparency, Consent, and Control)
On macOS, native messaging hosts can inherit browser TCC sandbox restrictions. After browser restarts, writes to protected paths (~/Documents, iCloud Drive, etc.) can fail with PermissionError.
Solution: The HTTP daemon (xtap_daemon.py) runs via launchd, independent of the browser process tree. It has its own TCC entitlements and can write to protected paths after a one-time macOS permission prompt. The daemon is the sole data transport — if it's not running, the extension buffers tweets and shows an error.
Tombstone tweets
X sometimes returns TimelineTweet entries where tweet_results.result is missing (deleted/suspended tweets). These are skipped by the parser. Since no ID is extracted, they don't enter seenIds — if the tweet later appears with full data, it will be captured.
Development Notes
- No build step — plain JS, no bundler, no transpilation. Load and go.
- Testing:
python3 -m pytest tests/ -v && node --test tests/*.test.mjs. Run after every change. CI runs these on every push to main with coverage uploaded to Codecov. For manual browser testing, load unpacked in Chrome (chrome://extensions) or as a temporary add-on in Firefox (about:debugging#/runtime/this-firefox) usingmanifest.firefox.json. Chrome extension IDs vary per install; Firefox uses the fixed Gecko ID inmanifest.firefox.json. Use the matching installer mode (install.sh ... chrome|firefox,install.ps1 -Browser chrome|firefox). - Parser golden test (fast, use for iteration):
node --test tests/parser-golden.test.mjs— runsextractTweetsagainst all fixture scenarios intests/fixtures/sanitized/*/and compares to goldenexpected.jsonl. Sub-second, reports field-level diffs on failure. Use this as the inner loop when modifying the parser. - E2E test (slow, CI gate):
cd tests/e2e && npm test— launches Chromium + real extension + daemon + Fake X server. Auto-discovers all scenarios intests/fixtures/sanitized/. Full pipeline integration proof. - Fixtures: Parser fixture packs live under
tests/fixtures/. Keep raw captures intests/fixtures/private-raw/only (gitignored) and commit only sanitized packs fromtests/fixtures/sanitized/. The anonymization methodology and review checklist are documented intests/fixtures/FIXTURES.md. - Adding a fixture from a dump: Discovery mode auto-dumps the first response per endpoint per session to the output dir as
dump-{endpoint}-{timestamp}.json(in{ endpoint, data }envelope format). Feed directly to the sanitizer:node tests/fixtures/tools/sanitize.mjs <dump-file> <scenario-name>. This producestests/fixtures/sanitized/<scenario-name>/with fixture.json, expected.jsonl, and manifest.json. Both parser and E2E tests pick it up automatically. - Workflow for supporting a new endpoint: (1) Human enables discovery mode and browses X normally — dumps appear automatically for every endpoint, (2) run
node tests/fixtures/tools/sanitize.mjs <dump-file> <scenario-name>, (3) runnode --test tests/parser-golden.test.mjs— if it fails, the parser doesn't handle this endpoint yet, (4) fix the parser inlib/tweet-parser.js, (5) re-run parser test until green (ms feedback loop), (6) regenerate expected: re-run sanitize.mjs (overwrites). - Debugging: Enable "Debug logging to file" in the debug dashboard (popup → "Debug Dashboard"). Logs write to
debug-YYYY-MM-DD.login the output directory. Service worker console is visible in each browser's extension debugger. The debug dashboard also shows live capture events with accept/dedup/error status, transport health, and a parser sandbox for testingextractTweetsagainst raw JSON. - Dev mode seenIds: In dev mode,
seenIdsuseschrome.storage.sessionwhen available (volatile — clears on extension reload). If session storage APIs are unavailable, it safely falls back tochrome.storage.local. - tweet-parser.js is the most fragile file — it handles multiple GraphQL response shapes and X changes their API schema without notice. The recursive fallback (
findInstructionsRecursive) catches many new endpoint shapes automatically, but field-level changes to tweet objects will need manual updates tonormalizeTweet(). - Service worker module:
background.jsis loaded as an ES module ("type": "module"in manifest). It importstweet-parser.jsdirectly. - HTTP daemon:
xtap_daemon.pybinds127.0.0.1:17381. Auth token stored at~/.xtap/secret(mode 600). Managed by launchd (macOS:launchctl kickstart -k gui/$(id -u)/com.xtap.daemon), systemd (Linux:systemctl --user restart com.xtap.daemon), or Scheduled Task (Windows:Stop-ScheduledTask -TaskName xTapDaemon; Start-ScheduledTask -TaskName xTapDaemon). Logs: macOS/Windows at~/.xtap/daemon-stderr.log, Linux viajournalctl --user -u com.xtap.daemon. - Transport debugging: The popup shows connection status and auto-refreshes every 2s. When transport is unavailable, the popup and debug dashboard show an actionable error message. Service worker console logs transport selection at startup and reprobe attempts. Daemon startup diagnostics are always logged to
~/.xtap/daemon-stderr.log; setXTAP_LOG_LEVEL=debugfor request-level detail. - Release checklist: (1) Bump
manifest.jsonversion, (2) bumpnative-host/xtap_daemon.pyVERSION, (3) runnode scripts/build-firefox-manifest.jsto regeneratemanifest.firefox.json. The manifest test validates version parity, so CI will catch a forgotten regeneration. (4) If any new files were added to the extension or native-host, update the file list in.github/workflows/release.yml— the release zip uses an explicit list, not a glob, so new files will be silently missing from releases if not added. (5) Follow the Release procedure below to publish. - Firefox manifest:
manifest.firefox.jsonis generated frommanifest.json— never edit it directly. The generator script (scripts/build-firefox-manifest.js) swapsservice_worker→scriptsand adds Gecko metadata.
Release Procedure
Every release follows the same flow so notes stay consistent across tags. Notes are always hand-written from the squashed PR body + the diff — never let the workflow auto-generate them.
- After the version-bump PR merges, pull main:
git checkout main && git pull --ff-only. - Draft notes locally using the template below. Save to
/tmp/release-notes-vX.Y.Z.md. - Pre-create the GitHub release as a draft with those notes:
gh release create vX.Y.Z --draft --title vX.Y.Z --notes-file /tmp/release-notes-vX.Y.Z.md - Push the tag:
git push origin vX.Y.Z. - The Release workflow detects the existing draft and uploads
xtap-X.Y.Z.zipinto it (no notes overwritten). - Verify the release page: zip attached, notes render correctly. Click Publish release when ready.
- Delete the local notes file and any merged feature branches.
If the workflow ever runs against a tag with no pre-existing release, it falls
back to gh release create --generate-notes. That fallback is a safety net,
not the intended path. Always pre-create.
Release Notes Template
Required sections (always present, in this order):
## Highlights
<1–3 bullets in plain English. What changed for users. Concrete, no hype.>
## Upgrading
<One line. Either: "No installer re-run needed — `git pull`, restart daemon,
reload extension." Or: "Re-run `install.sh` (or `install.ps1`) because <reason>.">
Optional sections (include ONLY when content exists; never ship an empty header):
## What's new— feature details, link to the PR (#NN).## Fixes— user-visible bug fixes, one bullet each, link to PRs.## Hardening— security-relevant changes (path validation, allowlists, size caps, etc.).## Tuning— new env vars / config knobs, with defaults.## Stats— test count, coverage, e2e scenarios. Trust signal — keep it factual.
Voice rules (same as the rest of the project):
- Concrete, specific, terse. Name files, env vars, PR numbers.
- No AI vocabulary (no "comprehensive", "robust", "nuanced", etc.) and no marketing copy.
- Reference issues and PRs by number so users can dig in.
- If a section would only have a single trivial bullet, fold it into Highlights instead of giving it its own header.
- Imperative, active voice in bullets. "Block redirects" beats "Redirects are now blocked."
When working with the user during a release, the assistant drafts the notes,
shows them in chat for approval, then runs the gh release create --draft
command. The user reviews on GitHub and clicks Publish.
Contributing
- Keep it simple. No build tools, no frameworks, no dependencies beyond Python 3 stdlib.
- Run
python3 -m pytest tests/ -v && node --test tests/*.test.mjsbefore submitting changes. - Every change must maintain zero network footprint. This is the core promise.
- Stealth constraints are non-negotiable — review the list above before submitting changes.
- Update README.md and AGENTS.md after every relevant change — new features, changed behavior, new config options, output format changes, new endpoints, architectural changes, etc. Both files must stay in sync with the code.
