Scholar-Agent is an agentic retrieval-augmented generation system for answering questions over a local collection of academic PDFs. It decomposes a question into evidence requirements, retrieves and reranks page-grounded passages, checks whether each requirement has enough support, performs bounded recovery when evidence is missing, and writes an answer with validated physical-page citations.
The project is designed around a simple principle: retrieval is not treated as complete when relevant text is merely found; the system explicitly assesses whether each requirement has sufficient evidence before generation.
Question
↓
Planner
↓
Researcher
↓
Assessment-first Controller
├── sufficient ─────────────────────→ Writer
├── missing ─→ bounded Recovery ───→ Writer
└── unresolved ─────────────────────→ Writer
↓
Citation Validator
↓
Answer
LangGraph connects the Planner, Researcher, Controller, Recovery, Writer, and Citation Validator through one shared state. The state contains the original question, atomic requirements, selected evidence, the Requirement–Evidence Blackboard, Controller and recovery audit traces, and the final answer.
PDFs are extracted one physical page at a time with PyMuPDF. Chunks never cross page boundaries, so every passage retains stable source metadata:
{
"chunk_id": "...",
"paper": "paper.pdf",
"page": 4,
"chunk_index": 12,
"page_chunk_index": 1,
"text": "...",
}Scholar-Agent builds both a BM25 index and a dense-vector index. Each index stores a corpus fingerprint and refuses to load when it no longer matches the processed chunks.
The Planner turns a question into one to five atomic evidence requirements. Each requirement contains a focused query, relevant target names, a retrieval strategy, and a bounded retrieval depth:
{
"id": "R1",
"description": "Explain how Self-RAG controls retrieval.",
"targets": ["Self-RAG"],
"query": "Self-RAG reflection tokens retrieval control",
"retrieval_strategy": "hybrid", # bm25 | dense | hybrid
"top_k": 8,
}Planner output is sanitized before execution. Unknown strategies fall back to hybrid retrieval, depths are clamped to the supported range, empty queries are replaced, duplicates are removed, and malformed plans receive a safe fallback requirement.
The Researcher executes the route selected for each requirement:
bm25: BM25(query) ───────────────────────→ rerank
dense: Dense(query) ──────────────────────→ rerank
hybrid: BM25(query) ─┐
├→ reciprocal rank fusion → rerank
Dense(query) ┘
All routes use the same cross-encoder reranker. Candidate selection reserves local coverage for individual requirements before filling the shared candidate pool. Evidence allocation balances requirement coverage, named targets, source diversity, and physical-page diversity.
Selected passages receive stable E1, E2, ... identifiers and are organized
on a Blackboard:
evidence_board = {
"R1": {
"requirement": "Explain Self-RAG retrieval control.",
"evidence_ids": ["E1", "E3"],
"candidate_papers": [...],
"status": "sufficient",
"covered": ["retrieval control", "reflection tokens"],
"missing": [],
"action": None,
},
"R2": {
"requirement": "Explain Corrective RAG document repair.",
"evidence_ids": ["E2"],
"candidate_papers": [...],
"status": "missing",
"covered": ["document refinement is identified"],
"missing": ["the decompose-filter-recompose procedure"],
"action": {
"tool": "search_within_paper",
"candidate_id": "P1",
"paper": "CRAG.pdf",
"query": "decompose filter recompose document refinement",
"state": "pending",
},
},
}One passage may support multiple requirements. Requirements with no selected
support remain visible with an empty evidence list instead of disappearing.
The Researcher initializes status to unknown; the Controller then writes
its coverage assessment and any sanitized recovery action directly into the
Blackboard. Evidence items retain supports as the reverse evidence-to-
requirement index and requirement_scores as the underlying relevance data.
The Controller inspects the first complete Researcher observation once per question. It returns one assessment for every planned requirement:
{
"requirement_id": "R2",
"status": "missing", # sufficient | missing | unresolved
"covered": ["document refinement is identified"],
"missing": ["the decompose-filter-recompose procedure"],
"action": {
"tool": "search_within_paper",
"candidate_id": "P1",
"query": "decompose filter recompose document refinement",
},
}sufficientmeans all requested aspects have direct support.missingmeans a displayed bounded action can seek the absent evidence.unresolvedmeans evidence is missing and no valid bounded action is available.
Every sanitized assessment is written to its requirement's Blackboard entry. The Controller trace keeps accepted actions and rejection metadata as an audit log, while workflow routing and Recovery read executable actions from the Blackboard. Recovery is limited to one round and at most two actions:
search_within_papersearches a paper already identified by the Researcher;expand_neighborsinspects adjacent chunks around selected evidence;increase_depthcombines the original requirement with a new Controller-generated query, raisestop_ktoMAX_TOP_K, and reruns the requirement's full retrieval route.
Actions use stable candidate-paper selectors and chunk IDs. Retrieved
candidates still pass through reranking and evidence selection before they can
enter the Writer context. Recovery refreshes evidence_ids without discarding
the Controller's requirement-level assessment and marks executed Blackboard
actions so they cannot be selected twice.
The Writer receives the original question, Controller coverage assessments,
and the full candidate evidence pool. Requirements remain research scaffolding
rather than an answer outline: the Writer selects only the evidence needed for
the shortest sufficient answer. It may use only supplied passages, must cite
factual claims with evidence IDs such as [E1], and mentions an evidence gap
only when it blocks an important part of the user's request. An empty evidence
set produces a deterministic abstention without a Writer call.
The Citation Validator replaces evidence IDs with physical-page citations:
[E1] → [Self-RAG.pdf p.1]
Unknown evidence IDs and fabricated page references are removed. This gives every retained citation a deterministic provenance path from answer to chunk to physical PDF page.
The two most recent evaluations use the updated 50-question academic-RAG
benchmark, deepseek-v4-flash at temperature zero, alternating paired Writer
order, frozen runtime hashes, and variant-blind review with cross-review and
adjudication before unblinding.
The experiments answer different questions and use independently generated paired runs, so their absolute percentages should not be compared across rows.
This experiment compares the complete architecture with a Simple RAG pipeline:
Original question
→ BM25 top-8 + Dense top-8
→ reciprocal rank fusion
→ cross-encoder rerank
→ fixed evidence budget
→ Writer
→ Citation Validator
| Quality metric | Simple RAG | Scholar-Agent | Improvement |
|---|---|---|---|
| Strict Success | 74.0% (37/50) | 94.0% (47/50) | +20.0 pp |
| Requirement Accuracy | 80.3% (57/71) | 98.6% (70/71) | +18.3 pp |
| Initial Retrieval Recall | 49.0% | 76.5% | +27.5 pp |
| Selected Evidence Recall | 35.3% | 62.7% | +27.5 pp |
The paired run produced 12 Strict Success repairs and 2 regressions. The exact
two-sided McNemar p-value is 0.0129. Requirement outcomes produced 13 repairs
and no regressions, with p=0.000244. Both improvements are statistically
significant within this benchmark.
Trace inspection linked four repairs to the complete Controller recovery chain:
missing assessment
→ recovery action
→ new evidence
→ changed Writer context
→ repaired answer
The remaining gains came from atomic requirement decomposition, targeted initial retrieval, requirement-aware evidence selection, and Blackboard-guided writing. The result supports the architecture as a whole rather than assigning all improvement to a single component.
The second experiment freezes the same initial Planner output, retrieval results, rerank candidates, selected evidence, and Blackboard for each pair. One variant writes directly from that state; the other runs the assessment-first Controller and bounded Recovery before using the same Writer policy.
| Quality metric | Without Controller | With Controller | Observed improvement |
|---|---|---|---|
| Strict Success | 88.0% (44/50) | 96.0% (48/50) | +8.0 pp |
| Requirement Accuracy | 94.4% (67/71) | 97.2% (69/71) | +2.8 pp |
| Citation Support | 99.7% (396/397) | 100.0% (467/467) | +0.3 pp |
The Controller produced four Strict Success repairs and no regressions. Its
exact McNemar p-value was 0.125, so this isolated 50-question result is a
positive paired signal rather than a statistically proven standalone gain.
Manual attribution found two requirement repairs that followed the full
assessment-to-recovery-to-answer chain.
Taken together, the experiments support the current design: the complete Scholar-Agent pipeline significantly improves over Simple RAG, and removing the Controller weakens the same system in the paired ablation. The architecture is therefore justified by both end-to-end quality and component-level evidence.
Requirements:
- Python 3.11 or newer
- uv
- a DeepSeek or OpenAI-compatible API key
Install the project:
git clone <repository-url>
cd scholar-agent
uv syncConfigure the current architecture:
export DEEPSEEK_API_KEY=...
export SCHOLAR_AGENT_LLM_MODEL=deepseek-chat
export SCHOLAR_AGENT_RECOVERY_MODE=controllerAn OpenAI key can be used instead:
export OPENAI_API_KEY=...
export SCHOLAR_AGENT_LLM_MODEL=gpt-4.1-miniIngest a directory of PDFs, build the indexes, and ask a question:
uv run scholar-agent ingest path/to/papers
uv run scholar-agent index
uv run scholar-agent ask "Compare how Self-RAG and Corrective RAG handle retrieval quality."The same workflow is available from Python:
from scholar_agent.config import Settings
from scholar_agent.llm import LLMClient
from scholar_agent.retrieval import RetrievalEngine
from scholar_agent.workflow import run_question
settings = Settings.from_env()
engine = RetrievalEngine.load(settings)
llm = LLMClient.from_env(settings)
state = run_question(
"How does Sentence-BERT make semantic search efficient?",
engine,
settings,
llm,
)
print(state["answer"])Important environment variables:
| Variable | Purpose | Example |
|---|---|---|
DEEPSEEK_API_KEY |
DeepSeek API authentication | sk-... |
OPENAI_API_KEY |
OpenAI API authentication | sk-... |
SCHOLAR_AGENT_LLM_MODEL |
Planner, Controller, and Writer model | deepseek-chat |
SCHOLAR_AGENT_EMBEDDING_MODEL |
Dense retrieval model | sentence-transformers/all-MiniLM-L6-v2 |
SCHOLAR_AGENT_RERANKER_MODEL |
Cross-encoder reranker | cross-encoder/ms-marco-MiniLM-L6-v2 |
SCHOLAR_AGENT_MIN_RERANK_SCORE |
Evidence retention threshold | -1.0 |
SCHOLAR_AGENT_DATA_DIR |
Processed corpus and index directory | data |
Run the quality suite with:
make quality- Scholar-Agent answers from the indexed local PDF collection; it does not automatically search the open web for missing sources.
- Retrieval quality is limited by corpus coverage, PDF extraction quality, embedding quality, and cross-encoder ranking.
- The Controller can misjudge evidence sufficiency or choose an unproductive recovery action. Its decisions remain bounded and fully traceable, but they are not guaranteed to repair every evidence gap.
- Citation validation proves that a citation maps to a retrieved physical page; it does not perform semantic entailment checking for every claim.
- Page-level gold annotations in the evaluation benchmark are non-exhaustive diagnostic signals, not complete definitions of evidence sufficiency.
- The reported experiments contain 50 questions from one academic corpus. Broader generalization requires larger and independently sampled benchmarks.
- Indexes are rebuilt as a unit rather than updated incrementally, and the current NumPy-backed dense index is intended for laptop-scale collections.
- Embedding and reranker models may need to be downloaded on first use.
MIT