Back to Home
Completed
3 Months

Agentic Research Workspace.

Enterprise agentic RAG knowledge base — upload your documents, ask questions, get citation-backed answers. Single Research Agent over hybrid retrieval with a swappable Model Gateway.

ReAct AgentHybrid RRFCitation-at-Retrievalpgvector + FTSModel GatewayAG-UI SSEarq WorkersLogfire Traced
AgentLangGraph ReAct
RetrievalHybrid RRF
Vectorpgvector 1024d
GatewayLiteLLM
Queuearq + Redis
StreamAG-UI SSE

Overview

Single-tenant MVP that proves the core loop: ask a question about your own documents, get a grounded answer — every claim traceable to a retrieved chunk, progress streamed in a fixed stage vocabulary.

Key focus areas:

  • Citation-backed RAG: HyDE → dense pgvector (1024d) + PostgreSQL FTS → RRF → Cohere rerank-v3 → cite chunks at retrieval time
  • Agentic orchestration: LangGraph Research Agent (ReAct, ≤5 iters, 60k token / 60s caps) over a retrieve_evidence tool
  • Async ingestion: upload returns immediately; arq worker does extract → Chonkie chunk → Cohere embed → Postgres in background
  • Model Gateway: every LLM call via LiteLLM complete / complete_with_tools (Groq → OpenRouter fallback, swappable without touching agent code)
Interactive Topology

Multi-Agent RAG Orchestration & Knowledge Graph

Explore the end-to-end knowledge architecture: async ingestion pipeline, hybrid retrieval (pgvector + PostgreSQL FTS + Cohere Rerank), LangGraph ReAct agent loop, and LiteLLM gateway. (Solid = Shipped MVP, Gold/Dashed = Planned V1 Specs).

Click to interact & explore topology

Impact & Results

< 60sIngestion

10-page PDF upload → searchable (extract → chunk → embed → pgvector).

100%Citations

Every claim rendered as [n] grounded to a chunk with source_ref.

≤ 5 itersRun Budget

60s wall / 30s per-call / 60k token caps — no runaway loops.


Features

Citation-at-Retrieval

Citations attached when chunks are retrieved, not hallucinated after. Answer [n] markers resolve against persisted citations.

Hybrid Retrieval

HyDE query expansion → dense ANN + PostgreSQL FTS tsvector merged by RRF → Cohere rerank-v3 for relevance.

AG-UI Streaming

POST /api/research/chat opens SSE: RUN_STARTED → CUSTOM stages → CUSTOM citations → TEXT_MESSAGE_CONTENT → RUN_FINISHED.

Async Ingestion

POST /api/ingestion/upload returns document_id + job_id immediately. arq worker advances queued → extracting → chunking → embedding → done.

Model Gateway

ADR-003: LiteLLM thin interface. Groq primary, OpenRouter fallback. Swap models via env, zero agent changes.

Multi-source Extract

Docling for PDFs, faster-whisper for audio, gitingest+tree-sitter for repos. Chonkie semantic/code chunkers preserve structure.


Low-Level: Ingestion

Ingestion Pipeline — upload to searchable

Off the request path. Same container runs arq worker (max_jobs=4). Vectors are authoritative in Postgres — no sidecar index to sync.

Component Breakdown
Request path (sync)
├── POST /api/ingestion/upload → write B2 + insert documents/ingestion_jobs (queued) → enqueue arq → return {document_id, job_id}
 
Worker path (async, off HTTP)
├── arq Worker picks job (same Render container, max_jobs=4)
├── Extract by source_type (Docling / whisper / tree-sitter)
├── Chonkie SemanticChunker / CodeChunker → tsvector per chunk
├── Cohere embed-english-v3.0 batch (hash fallback only in dev)
└── INSERT chunks (text + vector(1024) + tsv GIN) → stage=done

Low-Level: Research (ask → cited answer)

Research Loop — AG-UI SSE + ReAct

Product path is SSE (TanStack AI useChat). Legacy WS retained for wscat. Fixed stage vocabulary per docs/ux.md.

Component Breakdown
Client (TanStack AI useChat)
├── POST /api/research/chat {question, threadId} — opens SSE (AG-UI)
└── Renders stages + [n] markers → CUSTOM citations → final answer markdown
 
Server — LangGraph StateGraph
├── entry → planner (cheap model, 1-3 queries, fallback [question]) → researcher ⇄ loop (≤5)
│   └── synthesize → END (per specs/research-pipeline.md)
├── Researcher ReAct: thought → retrieve_evidence → observation (bounded by 60k tokens / 60s / 30s per-call)
└── Budget caps: MAX_ITERATIONS=5, AGENT_TOKEN_BUDGET=60k, AGENT_RUN_TIMEOUT=60s
 
Retrieval
├── HyDE via Model Gateway → hybrid dense+FTS (same Postgres table, RRF) → Cohere rerank-v3
└── Workspace pre-filter — no cross-workspace leakage

Data Model

Data Model — Postgres is the source of truth

pgvector lives beside chunk text in one table. No second vector DB to host, sync, or rebuild (ADR-002 addendum).

Component Breakdown
Postgres (Neon) — authoritative
├── users / workspaces (workspace_id allowlist on every retrieval)
├── documents + ingestion_jobs (queued → extracting → chunking → embedding → done/failed)
├── chunks (text + embedding vector(1024) + tsv generated + GIN index + source_ref + document_id)
└── research_runs + citations (answer + [n] → chunk_id FK, persisted before streaming)

Key Architectural Decisions

Why pgvector in Postgres (not Pinecone/Turbovec)?

Single table holds chunks.text + embedding + tsv. No sync job, no rebuild. FTS ts_rank_cd + dense cosine merged by RRF in SQL. See ADR-002 addendum — Turbovec removed.

Why LangGraph?

Stateful StateGraph gives retrieve_evidence tool loop with checkpoint-ready shape (PostgresSaver planned, ADR-001). Adding Planner/Critic later is additive (nodes, not new agents).

Why LiteLLM Gateway?

ADR-003 thin interface complete / complete_with_tools. Groq primary (speed/free tier), OpenRouter fallback. Swap via env, no agent diff.

Why arq + Redis?

Ingestion is N4 async by spec — upload must return in ms. arq with rediss:// (Upstash TLS) keeps HTTP fast, worker in same Render container (max_jobs=4, free tier sleeps on idle).


Trade-offs

Pros

  • Single source of truth: one Postgres for metadata, vectors, FTS — ops simple
  • Grounded by construction: citation-at-retrieval prevents post-hoc fabrication
  • Bounded cost: iteration/token/time caps + fixed stage vocab → predictable UX & spend

Cons

  • Single-node agent: no Planner decomposition yet — multi-hop needs specs/research-pipeline.md planner node (cheap model, fallback [question])
  • No Critic loop: deferred per PRD non-goals until golden dataset shows need
  • Free-tier cold start: one Render service sleeps on idle; no keep-warm cron

Technical Challenges

Preventing Hallucinated Citations

The Problem

LLM can fabricate source refs if citations are generated after the answer.

The Solution

Attach citations at retrieval time — every chunk returned by retrieve() already carries chunk_id + source_ref. Agent drafts against that evidence, synthesize persists citations before streaming.

Result: 100% of [n] markers resolve to a real chunk_id FK; no invented refs.

Hybrid Retrieval Without a Sidecar Index

The Problem

Separate vector DB means sync lag, rebuild on schema change, and cross-store allowlisting.

The Solution

Store Cohere 1024d vectors in chunks.embedding (pgvector) + generated tsv + GIN in same Neon table. HyDE → dense ANN + tsv FTS → RRF → Cohere rerank-v3 in SQL.

Result: One query, one allowlist (workspace_id), no sync/rebuild path.

Keeping Uploads Responsive

The Problem

Extracting PDFs/repos/audio + embedding can take 30–60s — blocking HTTP would timeout.

The Solution

FastAPI writes B2 + inserts Document/IngestionJob (queued), enqueues arq job, returns immediately. Poll GET /api/ingestion/jobs/{id} for staged progress. Worker advances stages with Logfire spans.

Result: Upload returns in <200ms; 10-page PDF searchable in <60s end-to-end.

API Reference


Reliability & Observability

OptimizationResult
TracingLogfire spans per run
Rate Limit5 uploads / 20 chats per min
Timeouts30s per-call / 60s wall
Backfillscripts/backfill_embeddings.py (NULL vectors only, Cohere-guarded)
DeploymentSingle Render container: alembic → arq + uvicorn

Lessons Learned

  • Groundedness is a data contract, not a prompt trick — citations must ride with retrieval, not be asked for after.
  • One authoritative store beats two synced ones — pgvector + tsv in Postgres removed an entire sync/rebuild class.
  • Bound the agent before improving it — caps + fixed stage vocab made the MVP shippable; Critic/Planner wait on golden eval data (docs/specs/evaluation.md).
  • Gateway abstraction pays off — Groq → OpenRouter swap is env-only; agent code never branches on provider.

Screenshots

Liked this project?

Check out more of my work or get in touch.