Imported from Flexipie/oculum (
docs/AGENTS.md). Install upstream withnpx skills add Flexipie/oculum --skill docs. Copyright stays with the author.
Guide for AI Agents
This document helps AI agents (like Claude Code) quickly orient and contribute effectively to Oculum.
Quick Start (60 seconds)
What is Oculum?
AI-native security scanner for LLM-generated code. Uses a 3-layer detection architecture (pattern matching → structural heuristics → AI validation) to catch security issues with minimal false positives.
Key Tech: TypeScript monorepo, Next.js web app, OpenAI gpt-5-mini model currently used, Anthropic Claude Haiku potential backup, Supabase backend.
Delivery Surfaces: CLI, VS Code extension, GitHub Action, web dashboard.
Current Phase
Phase 5: Open Beta (see 1-current/status.md)
Active focus: User onboarding, feedback collection, marketing content, polish based on real usage.
How to Find Work
- Check Linear issues: Active tasks synced to 1-current/active-tasks.md
- Read current status: 1-current/status.md has sprint breakdown and priorities
- Use playbooks: 3-playbooks/ for step-by-step workflows
Agent Workflows
When Starting a New Task
1. Read: docs/1-current/status.md (current context)
2. Read: docs/1-current/active-tasks.md (avoid duplicates)
3. Read: Relevant playbook from docs/3-playbooks/
4. Read: CLAUDE.md in project root (project-specific rules)
5. Create todo list with TodoWrite
6. Execute task with frequent progress updates
7. Update: docs/1-current/recent-changes.md
When Implementing a Feature
Playbook: 3-playbooks/new-feature.md
Key principles:
- Read existing code first (never propose changes to code you haven't read)
- Run scanner test suite before and after changes
- Update relevant docs in
2-implementation/ - Add entry to
1-current/recent-changes.md
When Fixing a Bug
Playbook: 3-playbooks/fix-bug.md
Key steps:
- Reproduce the issue
- Find root cause (use Grep/Read tools, not assumptions)
- Write failing test
- Implement fix
- Verify test passes
- Document in recent-changes.md
When Refactoring
Playbook: 3-playbooks/refactor.md
Critical:
- Create snapshot tests BEFORE refactoring
- Refactor in small, reviewable chunks
- Run full test suite after each chunk
- Verify no behavior changes
- Document breaking changes
Project-Specific Context
Scanner Architecture
3 Layers:
-
Layer 1: Pattern matching (secrets, high-entropy strings, config issues)
- Files:
packages/scanner/src/layer1/*.ts - Fast, deterministic, regex-based
- Files:
-
Layer 2: Structural analysis (auth patterns, dangerous functions, AI-specific risks)
- Files:
packages/scanner/src/layer2/*.ts - Context-aware heuristics
- Files:
-
Layer 3: AI validation (Claude Haiku validates Layer 1+2 findings)
- File:
packages/scanner/src/layer3/anthropic.ts - High-context validation with full file + project context
- File:
Key files:
packages/scanner/src/index.ts- Main orchestratorpackages/scanner/src/types.ts- Core types, vulnerability categoriespackages/scanner/src/tiers.ts- Detector tier assignments
Adding detectors: 2-implementation/scanner/adding-detectors.md
Monorepo Structure
oculum/
├── apps/
│ └── web/ # Next.js dashboard (Supabase auth, API routes)
├── packages/
│ ├── scanner/ # Core scanner engine (main package)
│ ├── shared/ # Shared types (auth, scan contracts)
│ ├── cli/ # CLI tool (Commander.js, Ink TUI)
│ ├── vscode-extension/ # VS Code extension
│ └── github-action/ # GitHub Action
Important Patterns
Auth Contract:
- All clients use
@oculum/sharedtypes - Tier-based access:
local(free, local),verified(Pro, backend),deep(Enterprise, backend) - API key verification via
POST /api/v1/verify-key
Scan Modes:
full: Complete scan of all files (initial onboarding, deep audits)incremental: Changed files only (CI/CD, fast feedback, skips Layer 3)
False Positive Handling:
- Follow workflow in 2-implementation/scanner/fp-triage.md
- Add regression test to
__tests__/regression/known-false-positives.test.ts - Update AI validation prompt if needed (
layer3/anthropic.ts)
Common Scenarios
Scenario: "Add new vulnerability detector"
1. Read: docs/2-implementation/scanner/adding-detectors.md
2. Create detector file in packages/scanner/src/layer2/
3. Add detector to layer2/index.ts exports
4. Add VulnerabilityCategory to types.ts (if new category)
5. Add tier assignment to tiers.ts
6. Create benchmark fixture in __tests__/benchmark/fixtures/layer2/
7. Run: npm test (from packages/scanner/)
8. Update AI validation prompt if needed (layer3/anthropic.ts)
9. Document in docs/1-current/recent-changes.md
Scenario: "Reduce false positives"
1. Read: docs/2-implementation/scanner/fp-triage.md
2. Identify FP pattern (which detector, which context)
3. Add regression test to __tests__/regression/known-false-positives.test.ts
4. Implement fix in relevant detector (add context awareness)
5. Run: npm test (ensure regression test passes)
6. Re-run validation on test repos (validation-repos/)
7. Measure FP reduction
8. Document in docs/1-current/recent-changes.md
Scenario: "Refactor large file (e.g., anthropic.ts, dangerous-functions.ts)"
1. Read: docs/3-playbooks/refactor.md
2. Run current test suite (establish baseline)
3. Create snapshot of current behavior (if not exists)
4. Identify logical modules within the file
5. Extract one module at a time into separate file
6. Run tests after each extraction
7. Verify no behavior changes (snapshots should match)
8. Update imports in dependent files
9. Document in docs/1-current/recent-changes.md
Critical Rules
DO:
- ✅ Read files before modifying them
- ✅ Run tests before and after changes (
npm testfrom relevant package) - ✅ Update
docs/1-current/recent-changes.mdfor significant changes - ✅ Follow existing code patterns and conventions
- ✅ Use TodoWrite to track multi-step tasks
- ✅ Consult
/CLAUDE.mdfor project-specific guidance - ✅ Mark todos as
in_progressBEFORE starting work - ✅ Mark todos as
completedIMMEDIATELY after finishing
DON'T:
- ❌ Modify code you haven't read
- ❌ Skip tests or proceed with failing tests
- ❌ Create duplicate docs (check if doc already exists)
- ❌ Ignore TypeScript type errors
- ❌ Batch multiple unrelated changes together
- ❌ Create new files when editing existing ones would suffice
Knowledge Sources
Primary (Always Current)
- CLAUDE.md (project root) - Project instructions, architecture, scanner patterns
- docs/1-current/status.md - Current phase, sprint, active work
- docs/1-current/active-tasks.md - Linear task sync
- docs/1-current/recent-changes.md - Last 7-14 days of significant changes
Secondary (Reference)
- docs/0-project/ - Project context, architecture, requirements
- docs/2-implementation/ - Domain-specific implementation guides
- docs/3-playbooks/ - Step-by-step workflows
- docs/5-reference/ - API contracts, roadmap, environment setup
Tertiary (Historical)
- docs/4-history/ - Completed work, ADRs, archived fixes
Anti-Patterns to Avoid
❌ Working in a Vacuum
Don't: Start coding without reading current status and active tasks.
Instead: Always check docs/1-current/ before beginning work.
❌ Proposing Changes to Unread Code
Don't: Suggest changes without reading the actual implementation.
Instead: Use Read tool on all relevant files first, understand the patterns.
❌ Creating Redundant Docs
Don't: Create new documentation without checking if similar doc exists.
Instead: Search docs/ for existing content, update rather than duplicate.
❌ Ignoring Test Failures
Don't: Proceed with implementation when tests are failing.
Instead: Fix tests immediately or understand why they're failing before continuing.
❌ Batching Unrelated Changes
Don't: Mix multiple concerns (refactor + feature + bug fix) in one implementation.
Instead: One logical change at a time, separate PRs/commits.
❌ Over-Engineering
Don't: Add abstractions, features, or complexity beyond what was requested.
Instead: Keep solutions simple and focused on the immediate requirement.
File Naming Conventions
Documentation
- Use kebab-case:
adding-detectors.md,fix-bug.md,recent-changes.md - Date historical files:
2026-01-18-cli-fixes.md - Be specific:
jwt-false-positives.mdnotfixes.md
Code
- TypeScript files: camelCase for utilities, kebab-case for components
- Test files:
*.test.tsor*.spec.ts - Fixtures: Match the detector name (e.g.,
ai-rag-safety.tsfixture for detector)
Branches
- Feature:
feature/<short-description> - Fix:
fix/<issue-description> - Refactor:
refactor/<component-name>
Getting Help
If Stuck on Implementation
- Re-read CLAUDE.md (project-specific patterns and guidance)
- Search
docs/4-history/for similar past work - Check
docs/2-implementation/for domain-specific guides - Read related test files for usage examples
If Missing Context
- Use Grep to find relevant code patterns
- Use Task tool with Explore agent for codebase questions
- Read git history:
git log --oneline --follow <file> - Check Linear issues in
docs/1-current/active-tasks.md
If Uncertain About Approach
- Read relevant playbook in
docs/3-playbooks/ - Look for similar completed work in
docs/4-history/ - Use AskUserQuestion tool to clarify requirements
- Propose multiple options with trade-offs
Success Metrics
Good Agent Behavior
- ✅ Completes task without breaking tests
- ✅ Updates relevant documentation
- ✅ Follows existing code patterns
- ✅ Adds tests for new code
- ✅ Documents decisions in commit messages
Excellent Agent Behavior
- 🌟 Identifies and fixes related issues proactively
- 🌟 Improves code quality beyond immediate task
- 🌟 Updates outdated docs encountered during work
- 🌟 Reduces technical debt opportunistically
- 🌟 Adds helpful code comments for complex logic
Red Flags (Avoid)
- 🚫 Breaking existing tests
- 🚫 Creating code without tests
- 🚫 Ignoring TypeScript errors
- 🚫 Skipping documentation updates
- 🚫 Making assumptions instead of reading code
Scanner-Specific Guidance
Detector Development Rules
- Context awareness is critical - Most FPs come from lack of context
- Severity matters - Be conservative, downgrade when uncertain
- Test files, seed files, example code - Always downgrade to
infoseverity - Library code vs application code - Different risk profiles
- AI validation is expensive - Only validate categories that truly need it
Common FP Patterns to Watch
- JWT tokens flagged as AWS secrets (check for
eyJprefix) Math.random()in test/seed files (check file path context)- Localhost URLs in dev configs (check for
.env.local,config.dev.ts) - Static content in
innerHTML(check for string literals) - JSON.parse with try-catch (don't flag as missing error handling)
When to Update AI Validation Prompt
Update packages/scanner/src/layer3/anthropic.ts VALIDATION_PROMPT when:
- Adding new vulnerability category
- Discovering new FP pattern that AI should reject
- Changing severity criteria
- Adding new context detection rules
Structure of validation prompt:
- Section 1: General false positive patterns
- Section 2: Context-aware rules (test files, examples, libraries)
- Section 3: Category-specific rules (auth, BYOK, RAG, etc.)
Testing Strategy
Before Any Change
cd packages/scanner
npm test # Run all tests
After Detector Changes
npm test -- --testPathPattern=benchmark # Benchmark fixtures
npm test -- --testPathPattern=regression # Regression tests
After Validation Changes
cd validation-test
node run-validation.js --repo vercel/ai --depth verified
Before Committing
npm run lint # From project root
npm test # From relevant package
git diff # Review all changes
Emergency Procedures
Tests Failing After Your Changes
- Don't panic - this is expected during development
- Read the test failure message carefully
- If snapshot mismatch: Review if behavior change is intended
- If assertion failure: Fix the code or update the test
- Never commit with failing tests
Accidentally Broke Something
- Run
git diffto see all changes - Identify the breaking change
- Revert specific changes:
git checkout -- <file> - Or revert all changes:
git reset --hard HEAD - Start over with smaller, incremental changes
Can't Find Where to Make a Change
- Use Grep to search for relevant patterns
- Use Task tool with Explore agent: "Where is implemented?"
- Check git history:
git log --all --full-history --source -- <file> - Look at import statements to trace dependencies
Quick Reference
Most Commonly Modified Files
packages/scanner/src/layer2/*.ts- Detector logicpackages/scanner/src/layer3/anthropic.ts- AI validationpackages/scanner/src/types.ts- Type definitionspackages/cli/src/commands/*.ts- CLI commandsapps/web/src/app/api/v1/scan/route.ts- Backend scan endpoint
Most Commonly Read Files
CLAUDE.md- Project guidancedocs/1-current/status.md- Current prioritiespackages/scanner/src/index.ts- Scanner orchestrationpackages/scanner/src/__tests__/benchmark/fixtures/*- Test examples
Key Commands
# Run scanner tests
cd packages/scanner && npm test
# Run CLI locally
cd packages/cli && npm start scan /path/to/code
# Start web app
cd apps/web && npm run dev
# Lint all code
npm run lint
# Build all packages
npm run build
Remember
The goal is not perfect code, it's working code that solves the user's problem with minimal false positives and excellent UX.
Be pragmatic, test thoroughly, document clearly, and ask questions when uncertain.
