Overview
Single-tenant MVP that proves the core loop: ask a question about your own documents, get a grounded answer — every claim traceable to a retrieved chunk, progress streamed in a fixed stage vocabulary.
Key focus areas:
- Citation-backed RAG: HyDE → dense pgvector (1024d) + PostgreSQL FTS → RRF → Cohere rerank-v3 → cite chunks at retrieval time
- Agentic orchestration: LangGraph Research Agent (ReAct, ≤5 iters, 60k token / 60s caps) over a
retrieve_evidencetool - Async ingestion: upload returns immediately; arq worker does extract → Chonkie chunk → Cohere embed → Postgres in background
- Model Gateway: every LLM call via LiteLLM
complete / complete_with_tools(Groq → OpenRouter fallback, swappable without touching agent code)
Multi-Agent RAG Orchestration & Knowledge Graph
Explore the end-to-end knowledge architecture: async ingestion pipeline, hybrid retrieval (pgvector + PostgreSQL FTS + Cohere Rerank), LangGraph ReAct agent loop, and LiteLLM gateway. (Solid = Shipped MVP, Gold/Dashed = Planned V1 Specs).
Impact & Results
Features
Low-Level: Ingestion
Ingestion Pipeline — upload to searchable
Off the request path. Same container runs arq worker (max_jobs=4). Vectors are authoritative in Postgres — no sidecar index to sync.
Low-Level: Research (ask → cited answer)
Research Loop — AG-UI SSE + ReAct
Product path is SSE (TanStack AI useChat). Legacy WS retained for wscat. Fixed stage vocabulary per docs/ux.md.
Data Model
Data Model — Postgres is the source of truth
pgvector lives beside chunk text in one table. No second vector DB to host, sync, or rebuild (ADR-002 addendum).
Key Architectural Decisions
Why pgvector in Postgres (not Pinecone/Turbovec)?
Single table holds chunks.text + embedding + tsv. No sync job, no rebuild. FTS ts_rank_cd + dense cosine merged by RRF in SQL. See ADR-002 addendum — Turbovec removed.
Why LangGraph?
Stateful StateGraph gives retrieve_evidence tool loop with checkpoint-ready shape (PostgresSaver planned, ADR-001). Adding Planner/Critic later is additive (nodes, not new agents).
Why LiteLLM Gateway?
ADR-003 thin interface complete / complete_with_tools. Groq primary (speed/free tier), OpenRouter fallback. Swap via env, no agent diff.
Why arq + Redis?
Ingestion is N4 async by spec — upload must return in ms. arq with rediss:// (Upstash TLS) keeps HTTP fast, worker in same Render container (max_jobs=4, free tier sleeps on idle).
Trade-offs
Pros
- Single source of truth: one Postgres for metadata, vectors, FTS — ops simple
- Grounded by construction: citation-at-retrieval prevents post-hoc fabrication
- Bounded cost: iteration/token/time caps + fixed stage vocab → predictable UX & spend
Cons
- Single-node agent: no Planner decomposition yet — multi-hop needs
specs/research-pipeline.mdplanner node (cheap model, fallback[question]) - No Critic loop: deferred per PRD non-goals until golden dataset shows need
- Free-tier cold start: one Render service sleeps on idle; no keep-warm cron
Technical Challenges
API Reference
Reliability & Observability
Lessons Learned
- Groundedness is a data contract, not a prompt trick — citations must ride with retrieval, not be asked for after.
- One authoritative store beats two synced ones — pgvector + tsv in Postgres removed an entire sync/rebuild class.
- Bound the agent before improving it — caps + fixed stage vocab made the MVP shippable; Critic/Planner wait on golden eval data (docs/specs/evaluation.md).
- Gateway abstraction pays off — Groq → OpenRouter swap is env-only; agent code never branches on provider.
Screenshots
Liked this project?
Check out more of my work or get in touch.