Imported from elder234/cyber-threat-intel (
AGENTS.md). Install upstream withnpx skills add elder234/cyber-threat-intel. Copyright stays with the author.
AGENTS.md
Aegis CTI — modular, containerized cyber-threat-intelligence platform.
Repo map
services/api— Node.js API. Fastify 4 + TypeScript (ESM), REST + GraphQL (mercurius) + WebSocket, JWT/RBAC, Postgres viapg, Redis viaioredis. Entry:src/server.ts(auto-runs migrations + admin seed on boot).services/rust-core— Cargo workspace (crates/*). Bins:aegis-worker(consumescollectors/defaultjob queues),aegis-scanner(consumesscannerqueue + ad-hoc CLI, lib+bin),aegis-collectors,aegis-analyzer. Libs:aegis-common(config/db/jobs),aegis-ioc,aegis-osint,aegis-container,aegis-malware.web— the frontend (React 18 + Vite + TS + Tailwind). Dev server on :5173 proxies/apiand/wstolocalhost:8080.db/migrations/— ordered SQL migrations (0001–0012), the single schema source of truth; alsodb/migrate.sh.deploy/— Helm chart + raw k8s manifests.docs/DATABASE.mddocuments the schema/queue design. (README links todocs/ARCHITECTURE.md,SECURITY.md,LEGAL.md— those files do not exist yet.)
Commands
- API (
services/api):npm install && npm run dev(tsx watch, port 8080). Verify withnpm run build(tsc),npm test(vitest). Env comes from repo-root.env(copy.env.example). - Rust (
services/rust-core):cargo test --workspace,cargo clippy --workspace -- -D warnings,cargo fmt --check. Build a single binary withcargo build --release --bin aegis-worker(same foraegis-scanner,aegis-collectors,aegis-analyzer). - Web (
web):npm run typecheck,npm run lint,npm run build. - Full stack:
cp .env.example .env && docker compose up -d --build. Before first boot, wire up the Cloudflare Tunnel (deploy/cloudflared/config.yml.example, see README). No host ports are published — dashboard, API, and all datastores areexpose:-only, and the sole public entrypoint is thecloudflaredtunnel athttps://<yourdomain>(the VPS makes outbound connections only; hit API health viacurl https://<yourdomain>/api/health). Rust services build ONE shared image (ghcr.io/elder234/aegis-cti-rust) that compiles all four binaries in a single cargo invocation (BuildKit cache mounts); compose selects the binary per service viacommand: ["/usr/local/bin/aegis-<name>"]. Never build the four as separate images — it recompiles the dep tree 4× in parallel and OOMs small VPSes.
Gotchas
- Node lockfiles are NOT committed (no
package-lock.json). Usenpm install, nevernpm ci.Cargo.lock(services/rust-core) IS committed — generated with cargo 1.97; regenerate it whenever you change workspace deps (cargo update/cargo generate-lockfile). - API TS is ESM: relative imports use explicit
.jsextension (from './pool.js'). Match this or tsx/node resolution breaks. - Rust builds set
SQLX_OFFLINE=true(Dockerfile, CI) since there is no DB at build time. Crates use runtimesqlx::query()/query_as()only — do NOT add compile-timequery!()/query_as!()macros, which needcargo sqlx prepareagainst a live DB (seeservices/rust-core/.sqlx/README.md). - Adding a migration: drop
NNNN_name.sqlintodb/migrations/; bothdb/migrate.shand the API's startup migration runner pick it up automatically (idempotent viaaegis.schema_migrations). The API Dockerfile copiesdb/→/db. - All DB objects live in the
aegisschema (search_path=aegis). Theseverity/confidenceenum declaration order is load-bearing —upsert_iocmerges upward viaGREATEST()on enum ordinals. - Auth: short-lived JWT + opaque refresh token (sha256, rotated with family-wide revocation). Access token is mirrored to
window.__aegisAccessTokenfor the WS layer. Default admin seeded fromSEED_ADMIN_EMAIL/SEED_ADMIN_PASSWORD. - WebSocket endpoint:
GET /ws; the client authenticates with a first message{"type":"auth","token":"<jwt>"}(never in the URL/handshake headers). Fan-out from Rediseventschannel, filtered per client by the caller's permissions. - Job queue contract (API↔Rust):
feed.pull→collectorsqueue (handled byaegis-worker);scan.run→scannerqueue (handled byaegis-scanner). API enqueues viaaegis.*stored procedures.
Verification status
API (npm run build) and web (typecheck/build/lint) are verified green and API unit tests pass (43/43). Rust compiles and its gates are green locally: cargo test --workspace (155 tests), cargo clippy --workspace -- -D warnings, cargo fmt --check. The many ⚠️ RUNTIME VERIFICATION REQUIRED markers still stand — they flag network/socket/DB paths unit tests cannot exercise (sockets, TLS handshakes, Tor fetches, live DB writes); don't assume the full stack has run end-to-end.
Tests
- API: unit tests are hermetic (
npm testinservices/api).test/api.integration.test.tsis skipped unlessRUN_INTEGRATION=1with liveDATABASE_URL/REDIS_URL. - Rust: unit tests are pure logic (feed parsers, IOC normalization, port parsing, TLS/HTTP analysis):
cargo test --workspace. - CI (
.github/workflows/ci.yml): apibuild → test; rustcheck → test → clippy -D warnings; webtypecheck → build → lint; thendocker compose build.
Audit backlog (2026-07-31)
Findings from a full read-only review against a running stack. Every claim below was verified by reading the cited file — no speculation. Line numbers are from the commit at review time; re-grep if they have drifted.
Work these in the order given: P0 is remotely exploitable pre-auth, P1 boots insecure by default, P2 makes the dashboard lie, P3 is a data-integrity bug.
Resolution status
- P0 — PARTIAL→DONE. Datastore
ports:removed from docker-compose.yml (ac529bc), and compose now publishes zero host ports — the sole entrypoint is the Cloudflare Tunnel (cloudflared) athttps://<yourdomain>(web isexpose:-only). OpenSearch security was still disabled in the deploy paths; resolved in F1 below. - P1 — DONE.
config.tshard-fails on/^change_me|^ChangeMe123!$/whenNODE_ENV=production;SEED_ADMIN_PASSWORDhas no default. New migration0013addsdashboard:read. - P2 — DONE for the drifted endpoints, verified.
0013_align_dashboard_outputs.sqlmakes dashboard_stats emit exactly the 10DashboardStatskeys (diffed programmatically: 0 missing, 0 extra) incl. realrisk_score/by_severity/ingest_24haggregates; unified_search emits label/sub_label/rank/severity; v_attack_stats→tactic/count; v_top_sources→source/count/high_sev; v_recent_kev widened to theCveshape;alerts.ts:26selects explicit columns withbody AS summary.db/smoke_test.sqlasserts all 10 keys plusrisk_scoreint-ness andby_severitybucketing.SELECT *survives in 8 other places — see F3 below. - P3 — DONE, verified.
malware.tsusesawait data.toBuffer()+ adata.file.truncatedcheck → 413. - P4 — DONE, verified.
/api/health/readyreturns a bare status (detail logged viareq.log.warn); all four dashboard routes + GraphQLdashboardStatsrequiredashboard:read, and0013:118-126seeds the permission and grants it to admin/analyst/viewer so the new guard does not lock everyone out; the WS hub authenticates via first message, filters per event throughEVENT_PERM, and times out unauth'd sockets. - P5 — PARTIAL.
mem_limitis on every compose service and the README quick-start no longer instructs--build— both verified. But the limits sum to more RAM than the box has — see F2 below.
Follow-up backlog (2026-07-31, post-remediation review)
Verified against the tree after fd9334c/ac529bc. Three items from the first pass are
still open. Same rules as before: re-grep before changing, never edit an applied migration.
F1 — OpenSearch still runs with authentication disabled (was P0)
plugins.security.disabled=true is unchanged in all three deploy paths:
| File | Line |
|---|---|
docker-compose.yml |
71 — comment still reads # dev only; enable TLS+auth in prod |
deploy/k8s/opensearch.yaml |
56 |
deploy/helm/aegis/templates/opensearch.yaml |
54 |
Closing the host port shrank the blast radius from "unauthenticated cluster on the public
internet" to "unauthenticated cluster reachable by any container on the aegis bridge
network" — a real reduction, but not a fix. The k8s and Helm manifests are the
prod-targeted ones and they do the opposite of what the compose comment advises.
OPENSEARCH_INITIAL_ADMIN_PASSWORD (docker-compose.yml:70) stays inert until the plugin
is on.
Task: enable the security plugin in the k8s and Helm manifests at minimum, wire the admin credential through, and point the API's OpenSearch client at it with TLS. If compose stays insecure for local dev, make that explicit — rename the comment to say the k8s/Helm paths differ, so nobody ships the dev posture by copying it.
F2 — Compose memory limits over-commit the 4 GB box
Every service got a mem_limit, but they sum to 4224 MiB on a 4096 MiB host:
postgres 384 + redis 256 + opensearch 1024 + api 512 + collectors 512
+ worker 384 + scanner 512 + analyzer 256 + web 256 + tor 128 = 4224 MiB
That leaves nothing for the kernel, dockerd, sshd, or page cache. mem_limit is a ceiling
rather than a reservation, so this will not OOM under light load — but there is no headroom
by construction, and OpenSearch at 1g with bootstrap.memory_lock=true
(docker-compose.yml:68) plus memlock: -1 (:73) holds its heap resident, so that GiB
is not opportunistic.
Task: get the total under ~3.2 GiB. Cheapest path is OpenSearch 1g → 512m (drop
-Xms/-Xmx to 256m to match) and trimming the three Rust workers, which are queue
consumers and mostly idle. Verify with docker stats under load rather than by guessing.
Alternative, if the current sizing is genuinely needed: provision 2 GB of swap and say so
in the README so the requirement is not invisible.
F3 — SELECT * remains in 8 query sites
P2's regression guard was applied to the endpoints that had drifted, but not to the rest. Remaining:
| File | Lines |
|---|---|
services/api/src/routes/cves.ts |
54 |
services/api/src/routes/iocs.ts |
72, 98 |
services/api/src/routes/rules.ts |
90 |
services/api/src/routes/scans.ts |
31, 34, 35, 36 |
These are single-row-by-id reads and scan-detail fan-outs, so the risk is lower than the
dashboard aggregates that actually broke — but it is the same coupling that produced P2:
the API contract becomes whatever the table happens to contain, and adding a column to
iocs or scans silently changes the response shape. TypeScript will not catch it.
Task: replace each with an explicit column list matching the declared response type. Do this while the P2 context is fresh — it is mechanical now and archaeology later.
Follow-up resolution status (2026-07-31)
All three follow-up items are landed on main.
- F1 — DONE. Security plugin is ENABLED in
deploy/k8s/opensearch.yamland the Helm chart (values.yamlopensearch.securityDisabled: false); the admin credential is bootstrapped fromOPENSEARCH_INITIAL_ADMIN_PASSWORDwired from theaegis-secretssecret / Helm secret in both paths. Probes switched fromhttpGet /_cluster/healthtoexeccurl -kfs -u admin:${OPENSEARCH_INITIAL_ADMIN_PASSWORD} https://localhost:9200/...(the plugin serves TLS on 9200 with its generated certs).OPENSEARCH_NODEconfigmaps now point athttps://. Verified caveat: no API component currently talks to OpenSearch — theOPENSEARCH_*env vars are reserved and unconsumed, so there was no client to point at TLS. On the 1.9 GiB VPS (2026-08-01), OpenSearch 2.13 could not boot inside a 512m limit even with security + PA disabled (~1g RSS floor with its bundled plugin set), so the compose service is now profiled out of the default stack (docker compose --profile opensearch up -dto opt in on a box with RAM to spare; sized 1g/512m heap) and theapi→opensearchhealth dependency was removed — nothing uses OpenSearch, and it was blocking the whole stack. The k8s/Helm paths still run it secured. - F2 — DONE. Compose limits total 3200 MiB (was 4224), then OpenSearch was
profiled out of the default stack (see F1), leaving the default set at 2624 MiB of
ceilings (384 postgres + 256 redis + 512 api + 256 worker + 256 collectors + 384 scanner
- 256 web + 256 analyzer + 64 cloudflared). OpenSearch 512m/256m heap; worker 512m→256m; collectors 384m→256m;
scanner 512m→384m. Redis
--maxmemory512mb→192mb so it evicts inside its 256mmem_limitinstead of being OOM-killed.SCANNER_MAX_CONCURRENCYlowered to 64 in.env.example, k8s configmap, and Helm values. Note the audit's "4096 MiB host" was wrong for the VPS — it has 1.9 GiB and no swap, so ceilings are only safe because actual usage is far lower. Verify the sizing withdocker statsunder load.
- 256 web + 256 analyzer + 64 cloudflared). OpenSearch 512m/256m heap; worker 512m→256m; collectors 384m→256m;
scanner 512m→384m. Redis
- F3 — DONE. All 8
SELECT *sites replaced with explicit column lists matching the declared response types:cves.ts:54,iocs.ts:72/98,rules.ts:90,scans.ts:31/34/35/36(scan detail columns enumeratescan_ports/scan_tls/findings). RemainingSELECT *ingraphql/index.tsis pinned bymapCve/mapIoc, andalerts/engine.tscallsraise_alert()(a function row, not a table) — both out of scope.
Original findings (2026-07-31, first pass)
Kept for context and for the file:line evidence behind each fix. Read the resolution status above first — most of this is already landed. Only F1/F2/F3 above still need work.
P0 — Datastores are published to 0.0.0.0
docker-compose.yml publishes five host ports; four should not be reachable from the
internet on a public VPS:
| Service | Line | Mapping |
|---|---|---|
| postgres | 32-33 | 5432:5432 |
| redis | 53-54 | 6379:6379 |
| opensearch | 77-78 | 9200:9200 |
| tor (SOCKS) | 204-205 | 9050:9050 |
| web | 171-174 | 8080:8080 — intended public entrypoint, keep |
Docker's short "HOST:CONTAINER" syntax with no IP prefix binds all interfaces. None of
these need a host port: api reaches all three datastores by service name over the
internal aegis bridge network (api itself is correctly expose:-only, line 102-103).
Task: delete the ports: entries for postgres/redis/opensearch/tor, or prefix each
with 127.0.0.1: if host access is wanted for debugging.
Do not rely on a host firewall here. Docker inserts DNAT rules into the DOCKER
iptables chain, which bypasses a ufw/nftables INPUT policy. Closing the port in
compose is the actual fix.
Related: OpenSearch runs with the security plugin off — plugins.security.disabled=true
(docker-compose.yml:71, deploy/k8s/opensearch.yaml:56-57,
deploy/helm/aegis/values.yaml:59). With 9200 public that is unauthenticated
read/write/delete on the whole cluster. OPENSEARCH_INITIAL_ADMIN_PASSWORD (line 70) is
inert while the plugin is disabled. Enable security for any non-dev deploy, and treat it
as required even after the port is closed.
k8s/Helm are not affected by the port issue — no NodePort/LoadBalancer anywhere under
deploy/, all Services are ClusterIP.
P1 — Placeholder secrets pass validation and boot silently
services/api/src/config.ts validates secret shape but never value:
- Line 25-26:
JWT_ACCESS_SECRET/JWT_REFRESH_SECRETarez.string().min(16).change_me_access_secretis 23 chars andchange_me_refresh_secretis 24 — both pass and the app starts. A publicly-known signing key means forgeable admin tokens. - Line 34:
SEED_ADMIN_PASSWORD: z.string().min(8).default('ChangeMe123!')— a hardcoded default inside application code. Omit the env var entirely and the API seeds an admin atadmin@aegis.localwith a password published in this repo, with no warning. docker-compose.yml:28,47,70use${VAR:-change_me_*}shell defaults, so the datastores also start with placeholder passwords when.envis missing.- Nothing checks
NODE_ENV === 'production';NODE_ENVitself defaults todevelopment(config.ts:12).
Net effect: git clone && docker compose up -d --build with no .env edit yields public
datastores + known JWT signing key + known admin login. CI does cp .env.example .env, so
the placeholder set is a tested, working configuration.
Task: add a zod .refine() (or a boot-time assertion) that hard-fails when any secret
matches /^change_me|^ChangeMe123!$/ while NODE_ENV === 'production'. Drop the
.default('ChangeMe123!') at line 34 outright — a seed admin password should have no
default. Consider raising the min(16) floor and rejecting low-entropy values.
Secrets hygiene is otherwise clean and should stay that way: no real credentials are
committed, .gitignore:1-4 correctly excludes .env with a !.env.example negation, and
a full-history scan shows .env was never committed.
P2 — Frontend types do not match what the DB returns (6 endpoints)
Root cause: types in web/src/lib/types.ts were transcribed from route handlers rather
than from live responses (the file header at lines 1-4 says so), and the drifted routes
are exactly the ones using SELECT * — nothing pins columns to types. TypeScript cannot
catch this; the mismatch is between SQL output and a hand-written interface.
Decide the fix direction once and apply it consistently. Renaming the SQL to match the
frontend is usually right (the UI names are better), but it means a new migration —
never edit an applied migration in place, add 0013_*.sql. Renaming the frontend instead
is cheaper but propagates DB naming into the UI.
P2.1 — dashboard_stats() — 6 of 10 fields missing (this is what the UI shows)
aegis.dashboard_stats() (db/migrations/0007_procedures.sql:188-201) returns 10 keys;
DashboardStats (web/src/lib/types.ts:108-119) declares 10; only 4 overlap
(iocs_active, cves_kev, alerts_open, scans_running).
Frontend expects but DB never returns: iocs_total, feeds_healthy, feeds_total,
risk_score, by_severity, ingest_24h.
DB returns but frontend ignores: iocs_critical, cves_total, alerts_critical,
findings_open, feeds_enabled, jobs_pending.
Visible symptoms: Feeds healthy renders undefined/undefined (Dashboard.tsx:94),
Ingest 24h blank (:107), risk gauge always 0/100 because Dashboard.tsx:56 falls
back to ?? 0, and the severity panel always shows the "No active indicators" empty state
(:151) even with 15k+ indicators loaded. A SOC dashboard that always reports zero risk
is worse than one that errors — it looks authoritative while conveying nothing.
Note feeds_enabled is the intended source for feeds_healthy/feeds_total, so part of
this is rename-drift rather than genuinely absent data. risk_score, by_severity, and
ingest_24h have no source at all and need real aggregate SQL written.
Also fix the test that missed this: db/smoke_test.sql:57-59 asserts only that
iocs_active and jobs_pending exist — both in the matching set. Assert the full key set
the frontend consumes.
P2.2 — v_attack_stats — 3 of 3 field names wrong
AttackStat (types.ts:128-132) declares tactic, technique, count. The view
(db/migrations/0008_views.sql:22) produces tactic_id, tactic_name, ioc_count —
zero overlap. technique has no source column anywhere; the view aggregates at tactic
level only, so decide whether to add it or drop it from the type.
Runtime effect, not just types: Dashboard.tsx:216 sorts on b.count - a.count → NaN
for every row; :226/:230 bind dataKey="tactic"/dataKey="count" → the ATT&CK chart
renders empty bars with blank axis labels.
P2.3 — unified_search — 3 of 6 field names wrong
SearchResult (types.ts:139-146) vs RETURNS TABLE at
db/migrations/0007_procedures.sql:159-161: label→title, sub_label→subtitle,
rank→score. severity is declared but never produced. entity_type/entity_id match.
Effect: Search.tsx:90 renders {r.label} → blank row text; :91 guards r.sub_label
→ subtitles never render; :89 guards r.severity → severity chips never render.
P2.4 — v_top_sources — count vs total
TopSource.count (types.ts:136) vs view column total (0008_views.sql:31). View also
returns high_sev (:32), undeclared in the type.
P2.5 — v_recent_kev — partial shape
cves.recentKev() (web/src/lib/api.ts:189) claims Cve[], but the view
(0008_views.sql:40-41) returns only cve_id, cvss_v31_score, epss_score,
kev_added_at, kev_ransomware, summary. Missing vs Cve (types.ts:57-67):
description (view aliases it to summary), cvss_v31_severity, epss_percentile,
kev, published_at. Either widen the view or give the endpoint its own narrower type.
P2.6 — Alert.summary has no matching column
Alerts.tsx:78 renders {alert.summary}; Alert.summary (types.ts:77) has no column
in the alerts table — it is body. alerts.ts:26 uses SELECT *, so the field is simply
absent at runtime.
Guard against regression: after fixing, replace SELECT * in alerts.ts:26,
search.ts:30, dashboard.ts:26, dashboard.ts:33, cves.ts:62 with explicit column
lists. The routes that already do this (iocs, rules, feeds, scans, channels,
alertRules, container, malware, auth.login, dashboard.timeline) were all
verified clean — explicit columns are why.
P3 — Malware upload silently truncates at 32 MiB and stores a wrong hash
services/api/src/routes/malware.ts:66-68 consumes the multipart stream by hand:
for await (const chunk of data.file) { chunks.push(chunk); }
@fastify/multipart has throwFileSizeLimit: true by default, but on overflow its
onError only stashes the error — it does not throw and does not destroy the stream. The
error surfaces solely via toBuffer() or the next parts iteration, and this route uses
neither. Busboy truncates at the limit and ends the stream normally, so the loop completes
without error and data.file.truncated is never checked.
Result: a >32 MiB upload is silently truncated to exactly 32 MiB, then hashed and
persisted as if complete. The sha256 written at malware.ts:112 (unique-indexed at
db/migrations/0012_malware_samples.sql:35) is the hash of a 32 MiB prefix, not of the
sample. For a threat-intel tool this is an evidentiary-integrity failure: the registry
asserts a hash matching no real artifact, and reputation lookups keyed on it will miss.
Task: use await data.toBuffer() (which re-throws RequestFileTooLargeError), or
check data.file.truncated after the loop and return 413. Add a test that posts >32 MiB
and asserts a 413 rather than a stored row.
Both malware.ts:16 and aegis-analyzer/src/main.rs:14 carry
⚠️ RUNTIME VERIFICATION REQUIRED markers — this is exactly the class of bug they predict.
Not a vulnerability, but worth knowing: uploads are buffered fully in memory
(malware.ts:64-69, forwarded at :78-82, never written to disk — no writeFile/
createWriteStream/saveRequestFiles anywhere in the module, so the "never touches disk"
design claim holds and the filename is not a traversal sink). On a 4 GB box concurrent
32 MiB uploads are a memory-pressure vector; RATE_LIMIT_MAX=300 is the only throttle.
There is also no MIME/extension validation — arguably correct for a malware endpoint, and
the route is gated by requirePerms('malware:run') (malware.ts:58).
P4 — Hardening gaps in an otherwise solid auth layer
State this plainly so nobody "fixes" what isn't broken: all 20 state-mutating routes
carry permission guards (real RBAC via requirePerms, not bare authentication), JWT
signature and expiry are genuinely verified (plugins/auth.ts:39 → req.jwtVerify();
no decode()-instead-of-verify anywhere), claims are loaded server-side from
aegis.user_permissions() (lib/auth.ts:40-62) rather than trusted from the token,
passwords are Argon2id with lockout after 5 failures, refresh tokens are stored as SHA-256
hashes and rotated with family-wide revocation on reuse detection
(routes/auth.ts:84-93), and config.ts:46-51 fails fast rather than falling back. Every
permission string used in a guard exists in the migrations — no typo'd guard silently
passing. Do not "simplify" any of this.
Remaining gaps:
/api/health/readyis public and leaks internals (routes/health.ts:15-32). Returns{ ok: false, error: err.message }per dependency (:27); Postgres/Redis driver errors carry internal hostnames, ports, and DB names. Latency values (:25) are an unauthenticated health oracle./api/healthbeing public is fine; this is a different route. Return a bare status + HTTP code, log detail server-side.- Dashboard routes check authentication but no permission
(
routes/dashboard.ts:8,15,24,31) — the only data reads using bareapp.authenticate. Any authenticated user gets aggregate counts plus 200 rows of threat timeline including titles (:18). Same gap in GraphQL atgraphql/index.ts:116-117, wheredashboardStatschecks onlyctx.userwhile every sibling resolver callsrequirePerm. Add adashboard:readcode. - WebSocket has no per-client filtering and takes the token in the query string
(
ws/hub.ts:25-48). Signature/expiry are verified correctly at:36, but any valid token receives the full Rediseventsbroadcast (:19-23) — aviewersees what an admin sees.?token=(:28) risks the JWT landing in proxy/access logs; noteapp.ts:41redacts theauthorizationheader but not query strings.
P5 — Stack does not fit 4 GB as configured
- No memory limits on any compose service. A grep for
mem_limit|deploy:|resourcesreturns only Redis's internal--maxmemory 512mb(docker-compose.yml:49-50), which is an eviction policy, not a container limit. Any container can consume all host RAM and trigger the OOM killer against an arbitrary victim. - Helm requests total exactly 3.0 GiB, limits 7.375 GiB (
values.yaml:22-158, countingreplicas: 2for api and web). On a 4 GB box that leaves ~1 GB for kernel, sshd, dockerd, and page cache. - OpenSearch heap is locked in RAM:
-Xms512m -Xmx512m(docker-compose.yml:69) withbootstrap.memory_lock=true(:68) andmemlock: -1(:73). Real RSS runs well above heap once metaspace, thread stacks, and Lucene mmap are counted. codegen-units = 1(services/rust-core/Cargo.toml:48-52) maximizes peak linker memory across 9 workspace crates — the most memory-hungry release setting available.
Mitigation already in place: every service names a prebuilt ghcr.io/elder234/aegis-cti-*
image, so plain docker compose up -d pulls instead of compiling. But
README.md:84 and the compose header both instruct --build, which is the worst case.
Task: add memory limits to every compose service; change the README quick-start to
omit --build for VPS deploys (build in CI, pull on the box). If building on-box is
unavoidable, add swap first and relax codegen-units. Also review the unconstrained
defaults SCANNER_MAX_CONCURRENCY=512, SCANNER_RATE_PPS=2000, WORKER_CONCURRENCY=8
(.env.example:54-56) for a shared 4 GB host.
Ground rules for whoever picks this up
- Verify before changing. Line numbers drift; re-grep the symbol rather than trusting the number. If a finding no longer reproduces, say so instead of "fixing" it blind.
- Never edit an applied migration. Schema changes go in a new
NNNN_*.sql. - Do not weaken the auth layer while touching adjacent code — see P4 for what is already correct.
- P2 needs one decision applied uniformly (rename SQL vs rename frontend). Do not fix half the endpoints one way and half the other.
- Rust remains unverified end-to-end (see Verification status above); a first compile may surface dependency drift unrelated to anything here.
Feature backlog (2026-07-31) — Dark-web monitor + deep web payload inspection
Two new capabilities. Both extend existing subsystems rather than adding new services, and both are safety-gated — read the Ground rules at the end of this section before writing any network code. Decisions already made by the operator (do not re-litigate):
- Web inspection depth: active DAST (benign, non-destructive probes), layered on top of a passive fingerprint pass that discovers where to probe.
- Dark-web sources: Tor-routed curated sources (leak/paste/forum pages), read-only.
- Target authorization: reuse the existing
aegis.assets.is_authorized = truegate — the same gateaegis-scanneralready enforces (crates/aegis-scanner/src/main.rs:101-112). No new authorization model.
Build order: F-DAST before F-DARKWEB — DAST reuses the scanner path you already know, darkweb touches Tor and needs the most safety review. Within each feature, do the DB migration and the passive/read-only stage first, wire it end-to-end to the UI, and only then add the active/expanding stage.
Existing anchors (verified — build on these, don't reinvent)
- Authorization gate:
aegis-scannerrefuses to scan a registered asset unlessassets.is_authorized = true(crates/aegis-scanner/src/main.rs:101-112,mark_scan_failedon refusal). Theassetstable isdb/migrations/0005_assets_scans.sql:7;kind IN ('host','domain','cidr','url','asn')already includesurl. - Scan/findings storage:
aegis.scans(0005:24,scan_typeis a free text tag — 'port','tls','http','subdomain','full'),aegis.findings(0005:88) already has exactly the columns DAST output needs:category,title,severity,cve_id(FK tocves),evidence jsonb,remediation,status. Reusefindings— do not add a parallel table. - Job queues: scanner consumes the
scannerqueue (aegis-scanner/src/main.rs:16const QUEUE = "scanner"); collectors consumecollectors. Enqueue viaaegis.enqueue_job($kind,$queue,$payload,...)(aegis-common/src/jobs.rs:79-87). The API enqueues scans fromservices/api/src/routes/scans.ts(gated byscan:run). - Tor: SOCKS proxy already in compose as
torservice behind--profile tor(docker-compose.yml:204), read from Rust viaconfig.tor_socks_proxy(aegis-common/src/config.rs:12,41, envTOR_SOCKS_PROXY). The dark-web collector must route every request through this proxy — never make a clearnet request to an onion or a leak site. - Next migration number is
0014(highest applied is0013_align_dashboard_outputs.sql). - Alerts: matches should raise alerts through the existing alert engine (Module 11) so they surface on the dashboard and notification channels, not a bespoke path.
- CVE correlation:
scan_ports.product/version/cpe(0005:52-56) already capture service versions; the CVE DB is synced (cvestable). Version→CVE matching is a stated goal of the passive pass.
F-DAST — Deep web payload inspection + vulnerable-port correlation
Extends aegis-scanner. Adds a web-application inspection scan type that (1) passively
fingerprints a target, (2) correlates discovered service versions against the CVE DB, and
(3) sends benign active DAST probes to discovered injection points. Active probes only run
against an asset with is_authorized = true — enforce the same gate the port scanner uses,
in the same place, before any probe traffic.
F-DAST.1 — Migration 0014_web_inspection.sql
- Add a scan_type value convention
'web'(scan_type is free text, no enum change needed; just document it). - New table
aegis.web_findingsonly iffindings.evidence jsonbproves insufficient — first attempt should map ontofindingswithcategory IN ('fingerprint','version_cve','xss','sqli','path_traversal','open_redirect','header','cookie')and structuredevidence(request, param, payload, response marker, confidence). Prefer reuse; adding a table is a fallback, and if you add one, add a new migration, never edit 0005. - Add a
web_probe_policyseed row or config documenting which probe classes are enabled and their payload catalog version, so a scan records what it was allowed to do.
F-DAST.2 — Passive fingerprint + version→CVE (do this first, no attack traffic)
New module crates/aegis-scanner/src/web/fingerprint.rs:
- Fetch the target over HTTP(S) using the existing client; capture status, headers, cookies, body. Detect server/framework/CMS from headers, cookie names, meta tags, and known paths (a small static signature table, not a network wordlist bust).
- Reuse
http_headers::analyze(already exists) for header findings. - Map detected
product+versiontocpe, then queryaegis.cvesfor matching CVEs and writefindingsrows withcategory='version_cve',cve_idset, severity from the CVE. - This stage is safe against any reachable host and is the discovery input for F-DAST.3.
F-DAST.3 — Active DAST probes (gated, benign, non-destructive)
New module crates/aegis-scanner/src/web/probes.rs:
- Gate first: assert
is_authorized = truefor the asset before sending a single probe; reuse the exact check ataegis-scanner/src/main.rs:101-112. Ad-hoc URL targets require the operator-asserts path AND thescan:runpermission, logged. - Probe classes, all non-destructive and idempotent:
- reflected-XSS canary (unique inert marker, detect verbatim reflection — never a live
<script>that executes anything server-side) - error-based SQLi marker (detect DB error signatures in the response; no stacked queries,
no
OR 1=1auth bypass against prod data, no time-based blind that hammers the server) - path traversal (bounded, e.g.
../to a known-safe canary path, detect signature) - open redirect (detect 3xx to an attacker-controlled marker host)
- reflected-XSS canary (unique inert marker, detect verbatim reflection — never a live
- Hard limits: honor
SCANNER_MAX_CONCURRENCY/SCANNER_RATE_PPS; never send destructive methods (no DELETE/PUT payloads, no form submissions that mutate state); cap payloads per endpoint; respect robots is NOT sufficient — the authorization gate is the control. - Write each confirmed issue as a
findingsrow withevidence= {method, url, param, payload, matched_marker, confidence}. Mark low-confidence asstatus='open'with a confidence field so the UI can separate confirmed from suspected.
F-DAST.4 — API + queue wiring
services/api/src/routes/scans.ts: acceptscan_type: 'web'with aprofiledescribing enabled probe classes; enqueue to thescannerqueue. Keep the existingscan:runguard and the asset-authorization check that the route already documents (scans.ts:15-16).- Add a
web:scanpermission if you want to separate web DAST from port scanning in RBAC (new migration, seed grant to admin/analyst). Otherwise reusescan:runand say so.
F-DAST.5 — Web UI
- Extend the existing Scans page (or add a Vulnerabilities/Web tab) to launch a
webscan against an authorized asset and renderfindingsgrouped by category with the evidence payload. ReuseDataTable/Panel/severity chips. Guard the numeric coercion trap from the post-remediation fix (Postgres NUMERIC → string) if any score fields render.
F-DARKWEB — Dark-web monitor (Tor-routed, read-only, watchlist-driven)
New collector module in aegis-collectors, polling curated public leak/paste/forum sources
exclusively over the Tor SOCKS proxy, matching page content against an operator watchlist
(brands, domains, email patterns), and raising alerts on hits. Read-only: fetch and parse
public pages, never authenticate, post, purchase, or interact.
F-DARKWEB.1 — Migration 0014b/0015_darkweb.sql
aegis.watchlist— id, kind (domain,email,keyword,brand,bin), value, label, severity, enabled, created_by, timestamps.aegis.darkweb_sources— id, name, kind (leak_site,paste,forum), url +is_onion(true= MUST route via Tor,false= clearnet indexer),format(html|json), enabled, last_polled_at, poll_interval, health.aegis.darkweb_hits— id, source_id, watchlist_id, url, snippet (redacted/truncated), matched_value, observed_at, severity, alert_id (FK once alerted), status. Unique on (source_id, url, matched_value) to dedupe re-observations.- Seed a small curated
darkweb_sourcesset of well-known ransomware leak indexes / paste sites. Do not hardcode illicit-market URLs that facilitate transactions — leak-site monitoring for victim/brand exposure is the scope; sourcing a shopping list is not.
F-DARKWEB.2 — Tor-routed collector
New module crates/aegis-collectors/src/darkweb.rs:
- Build the HTTP client to route through
config.tor_socks_proxy; fail closed — if the proxy is unset or unreachable, log and skip, never fall back to a direct connection (that would leak the platform's real IP and defeat the point). - Poll each enabled source on its interval, parse public listing/paste pages, extract text.
- Match extracted text against enabled
watchlistrows (exact domain/email, keyword contains, brand fuzzy). On match, upsertdarkweb_hits(dedup via the unique constraint). - Rate-limit and jitter requests; a monitor that hammers a hidden service is both rude and fingerprintable.
F-DARKWEB.3 — Alerting + API + UI
- On a new
darkweb_hitsrow, raise an alert through the existing alert engine (Module 11) so it hits the dashboard live feed and configured notification channels. Setalert_id. - API:
services/api/src/routes/— CRUD forwatchlist(permwatchlist:write) and read endpoints fordarkweb_hits/darkweb_sources(permdarkweb:read). New migration seeds the permissions and grants. - UI: a Dark-web tab listing hits (source, matched value, snippet, severity, observed_at) and a watchlist editor. Snippets must be truncated/redacted — do not render raw dumped credentials or PII in full in the console.
Ground rules for these two features (read before writing network code)
- The authorization gate is the control, and it is not optional. Active DAST touches only
is_authorized = trueassets, enforced in Rust before the first probe, mirroringaegis-scanner/src/main.rs:101-112. If you cannot see the gate in the code path you are writing, the probe does not run. - DAST payloads are benign and non-destructive. Detection markers only — no state mutation, no auth bypass against real data, no resource-exhaustion/time-based flooding, no destructive HTTP methods. This is exposure discovery, not exploitation.
- Dark-web fetching is read-only and Tor-only, fail-closed. No auth, no posting, no purchasing, no interaction with markets. Every request goes through the SOCKS proxy or does not happen. Scope is brand/victim/credential exposure on public leak/paste/forum pages.
- Redact on the way in. Truncate snippets, mask credential/PII payloads before they hit the DB and the UI. The platform stores evidence of exposure, not a usable copy of the dump.
- Legal/ethical scope is already stated in
README.md:8-11("publicly available data … only scan assets you are authorized to test … never bypass authentication or access controls"). These features must stay inside that statement; covered indocs/LEGAL.mdanddocs/SECURITY.md. - Rust build discipline unchanged:
SQLX_OFFLINE=true, runtimesqlx::query()only, no compile-time macros; regenerateCargo.lockif you add deps. New scan/collector logic should carry pure-logic unit tests (signature matching, payload/marker detection, watchlist matching) since the network paths can't run in CI. - New migration, never edit an applied one. Next free number is
0016.
Feature backlog v2 (2026-08-14) — pcap analysis, botnet exposure audit, video first-seen
Three features, build in this order. Step 0 must land first so the tree is clean.
Step 0 — Land pending dark-web clearnet work
Uncommitted: db/migrations/0016_darkweb_clearnet_sources.sql, darkweb.rs,
collectors main.rs, darkweb.ts, web types.ts, AGENTS.md/SECURITY.md edits.
Commit (exclude package-lock.json and opencode.json), push, CI builds, deploy via
git pull + docker compose pull && docker compose up -d --force-recreate on the VPS
(root@165.227.96.161, /root/cyber-threat-intel). Verify ransomware.live victims JSON
and the WikiLeaks V2 mirror poll clean.
F-PCAP — pcap/pcapng upload analyzer (migration 0017)
Upload a Wireshark capture; the analyzer returns a full report; only metadata is stored.
- Migration
0017_pcap_analysis.sql(mirror0012/0015style):aegis.pcap_analyses— id, sha256 UNIQUE, size_bytes, format, packet_count, duration_ms, interface_count, top_talkers jsonb, protocol_mix jsonb, dns_queries jsonb, tls_snis text[], http_hosts text[], suspicious jsonb, ioc_matches jsonb, score int, summary, requested_by, created_at/updated_at. No raw capture column.aegis.pcap_findings— id, analysis_id FK, finding_id, severity, title, detail (same shape asmalware_findings).- Permissions
pcap:read(admin/analyst/viewer),pcap:run(admin/analyst) + grants.
- New crate
crates/aegis-pcap(mirroraegis-malwarelayout; depspcap-parser0.17etherparse0.21 — both pure Rust, zero system deps, no libpcap/tshark):
parse.rscontainer (pcap+pcapng, streaming),packet.rs(etherparse SlicedPacket: Ethernet/VLAN/IPv4/v6/TCP/UDP/ICMP/ARP),flows.rs5-tuple aggregation → top talkers/ports/unique hosts,dns.rsqueries+counts (DGA signal),tls.rsSNI/version,http.rshost/URI/method from unfragmented packets,heuristics.rs(SYN-flood bursts, Mirai signature — SYN seq == dst IP, Bashlite!prefix, beacon-like regularity),ioc.rspure matcher over an IOC slice,lib.rsanalyze_capture(&[u8], &[Ioc]) -> PcapReport.- Analyzer stays offline: the API passes the active IOC list in the request body.
- Pure-logic unit tests craft synthetic pcap bytes; network paths carry
⚠️ RUNTIME VERIFICATION REQUIRED.
- Analyzer: add
/analyze/pcaptoaegis-analyzer/src/main.rsnext to/analyze(sameRequestBodyLimitLayer; a small JSON body envelope for the IOC list). - Dockerfile: UNCHANGED (pure Rust parsing).
- API
services/api/src/routes/pcaps.ts: POST/api/pcaps(stream → analyzer, 413 on overflow, upsert on sha256), GET/api/pcaps, GET/api/pcaps/:id, DELETE. Gatedpcap:run/pcap:read. IOC list pulled fromaegis.iocs(type IN ('ipv4','ipv6','domain','url')— enum from0001). - UI
web/src/pages/Pcaps.tsx+ nav +types.ts: upload dropzone, report card (packet count, duration, protocol bars, top talkers), findings table with severity chips, IOC-match panel linking to/iocs.
F-EXPOSURE-CORRELATION — botnet exposure audit (migration 0018)
Ties feeds/IOCs to assets: "does my asset talk to known-bad infrastructure, and is it running a service that makes it abusable as a proxy/relay (free bandwidth)?" Strictly defensive — probes authorized assets only, checks-but-never-uses credentials, output is a remediation list.
- Migration
0018_exposure_audit.sql:aegis.exposure_audits— id, asset_id FK, scan_id FK, proxy_http bool, proxy_socks bool, smtp_open_relay bool, dns_open_resolver bool, weak_auth_services jsonb (service/port/creds_checked, never the password), ioc_overlap jsonb (matched ioc ids/values), beacon_flows jsonb, score int, status, timestamps. Reusesaegis.scans(scan_type='proxy_audit') andaegis.findingsfor per-check detail rows — do NOT add a parallel findings table. Permissionsexposure:read(admin/analyst/viewer),exposure:run(admin/analyst). - Scanner
crates/aegis-scanner/src/exposure.rs— new scan category gated by the SAMEassets.is_authorizedcheck asmain.rs:101-112(mirror it before any probe):- SAFE probes: open HTTP/SOCKS proxy detection, SMTP relay HELO test, DNS recursion check, single non-destructive default-credential check on telnet/SSH (creds never used beyond one auth attempt).
- IOC cross-ref: asset's observed IPs/domains vs active
aegis.iocs— a host that both talks known-bad AND runs an abusable service → high finding "candidate for proxy/relay abuse". - Beacon detection: via F-PCAP analyzer on a capture attached to the asset (host beaconing to a C2 IOC = confirmed-compromised, critical).
- Pure-logic unit tests for probe-result classification and scoring.
- API
services/api/src/routes/exposure.ts: launch audit (reusescan:runor newexposure:run), GET audits, GET/:id, triage. UI: Exposure tab on Scans page ranking assets by abusable + known-bad overlap.
F-VIDEO — video fingerprinting + first-seen attribution (migration 0019)
Register a video once; store hashes + perceptual frame-set + ffprobe metadata ONLY (never bytes); monitor watched sources for near-duplicate sightings; build a first-seen timeline so the earliest appearance + device metadata narrows the original poster.
- Migration
0019_video_first_seen.sql:aegis.video_fingerprints— id, sha256 UNIQUE, md5, size_bytes, duration_ms, width, height, fps, codec, encoder, creation_time, gps, file_type, perceptual_frame_hashes text[], status, summary, requested_by, timestamps.aegis.video_sources— id, name, kind CHECK IN ('telegram','x','facebook','snapchat','manual'), config jsonb, enabled, poll_interval_secs, last_polled_at, health.aegis.video_sightings— id, fingerprint_id FK, source_id FK, url, author, snippet (redacted), observed_at, matched_via ('perceptual'|'exact'|'manual'), confidence numeric, alert_id FK, status; UNIQUE (source_id, url).- Permissions
video:read,video:run; alert-rule seedvideo.sighting.
- New crate
crates/aegis-video:ffprobe.rs(shellffprobe -print_format json→ typed metadata),frames.rs(ffmpeg rawvideo pipe, ~16-32 sampled frames scaled 160x90 gray),dhash.rs(9x8 difference hash → u64/frame),match.rs(Hamming ≤8, set-overlap → confidence),lib.rsanalyze_video(&[u8]) -> VideoReport. Pure-logic tests for dhash/match/ffprobe-JSON parsing; ffmpeg paths carry the RUNTIME marker. - Dockerfile: add
ffmpegto the runtime stage apt install (only system change). - Analyzer: add
/analyze/video;ANALYZER_MAX_BYTESenv, 128 MiB default for video. - Collector
aegis-collectors/src/video.rs+ main.rs wiring (mirrordarkweb.rs): poll duevideo_sources; Telegram connector is read-only Bot API (getUpdates) — the only autonomous source for now. X connector ships DISABLED until a paid API key exists. Facebook/Snapchat are manual-report-only via POST/sightings(no public API; never scrape ToS-protected endpoints). Jitter + rate-limit; redact snippets; fail per-source. - API
services/api/src/routes/video.ts: POST/fingerprints(stream → analyzer, 413, metadata-only), GET/fingerprints, GET/fingerprints/:id(+ sightings timeline), source CRUD, GET/POST/sightings, PATCH/sightings/:id,raiseAlertsForNewSightingsviaaegis.raise_alert(dedupe pattern fromdarkweb.ts:143). - UI
web/src/pages/VideoTracker.tsx+ nav +types.ts: upload card, fingerprint metadata grid, first-seen timeline, sightings table, source manager.
Ground rules for all three (do not skip)
- Metadata + hashes only — raw upload bytes NEVER persisted, logged, or forwarded (malware posture; analyzer drops bytes at end of request).
- New migration per feature, never edit an applied one. Runtime
sqlx::query()only (no compile-time macros); regenerateCargo.lockif deps change (F-PCAP addspcap-parser+etherparse; F-VIDEO adds none, uses ffmpeg). - Every API route
requirePerms-gated; do not weaken the auth layer. - Active probing only against
assets.is_authorized = true, mirrored in Rust before any probe; probes are non-destructive; credentials checked, never used beyond one attempt. - Snippets/evidence redacted on the way in; no ToS-bypass scraping of X/FB/Snapchat.
- Verify:
cargo test --workspace,cargo clippy --workspace -- -D warnings,cargo fmt --check; APInpm run build+npm test; webnpm run typecheck/npm run lint/npm run build.
