All work
02Agentic RAG Platform2026

CampusCopilot

A retrieval score only means something once the benchmark is clean and the assistant knows when to say nothing.

FastAPIPostgreSQLpgvectorDockerLLM tool-calling

Context & problem

Dense retrieval alone handles paraphrases well but misses exact names, course codes, and uncommon terms. A research assistant that answers confidently from documents that do not contain the answer is worse than one that says it does not know.

Architecture

Student question enters the pipeline.

The interesting engineering: A better score can still be a worse answer

Dense search handled paraphrases well; BM25 rescued exact names, course codes, and uncommon terms. Reciprocal rank fusion gave a practical way to combine both without pretending their raw scores were comparable.

The less glamorous work made the result trustworthy: removing leaked benchmark text, tying every answer to a page, testing multiple storage backends, and calibrating abstention. "I cannot find that in the documents" is often the most reliable answer a research assistant can give.

Result

88.2%
Hit@5

up from 66.4%, 110-question benchmark

0.68
MRR

up from 0.50

90.5%
out-of-scope abstention

106 tests across both storage backends

What I learned

  • Benchmark integrity is part of the engineering, not a formality. 28% of candidate questions were rejected for overlap or answer-leak before the score could be trusted.
  • Calibrating when a system should refuse to answer is as much design work as improving the score when it does answer.

Source & demo

No hosted demo yet. CampusCopilot expects a private document corpus, so the benchmark results and test suite in the repo are the fairest way to judge it.

View on GitHub

Next project

Conflux