CampusCopilot
A retrieval score only means something once the benchmark is clean and the assistant knows when to say nothing.
Context & problem
Dense retrieval alone handles paraphrases well but misses exact names, course codes, and uncommon terms. A research assistant that answers confidently from documents that do not contain the answer is worse than one that says it does not know.
Architecture
Student question enters the pipeline.
The interesting engineering: A better score can still be a worse answer
Dense search handled paraphrases well; BM25 rescued exact names, course codes, and uncommon terms. Reciprocal rank fusion gave a practical way to combine both without pretending their raw scores were comparable.
The less glamorous work made the result trustworthy: removing leaked benchmark text, tying every answer to a page, testing multiple storage backends, and calibrating abstention. "I cannot find that in the documents" is often the most reliable answer a research assistant can give.
Result
up from 66.4%, 110-question benchmark
up from 0.50
106 tests across both storage backends
What I learned
- Benchmark integrity is part of the engineering, not a formality. 28% of candidate questions were rejected for overlap or answer-leak before the score could be trusted.
- Calibrating when a system should refuse to answer is as much design work as improving the score when it does answer.
Source & demo
No hosted demo yet. CampusCopilot expects a private document corpus, so the benchmark results and test suite in the repo are the fairest way to judge it.
View on GitHubNext project
Conflux