NexusScholar
NexusScholar is VTG’s document-intelligence layer — a retrieval-augmented generation (RAG) capability that turns large PDF and document collections into a searchable knowledge base and answers natural-language questions with grounded, cited evidence. It combines semantic embedding, vector search and cross-encoder reranking with local large-language-model generation, and adds a prompt-driven toolkit to summarise, expand and reshape source text.
Built on the NexusFlow reasoning layer and assurable via NexusTrust, NexusScholar runs fully on-premise for data-sovereign deployments — no data leaves your infrastructure.
Key Features
- Semantic indexing of large PDF collections with a per-document embedding cache
- Retrieval with cross-encoder reranking for high-precision, grounded answers
- Page-level source citations on every answer for full traceability
- Prompt-driven toolkit — summarise, expand, simplify, abstract and rewrite
- Map-reduce summarisation of whole documents beyond the model’s context window
- Fully local, on-premise operation — no data leaves the deployment
Application Areas
- Research literature review and evidence synthesis
- Horizon Europe proposal and deliverable drafting support
- Regulatory and compliance document interrogation
- Technical-report and standards knowledge bases
- Clinical-guideline and protocol question answering
- Cultural-heritage and archival document exploration
- Public-sector policy and legal document analysis
- Internal knowledge management and staff onboarding
Technical Specification
| Architecture | Modular RAG pipeline: ingest → embed → index → retrieve → rerank → generate |
| Models | Local embedding + cross-encoder reranker with an Ollama-served LLM (swappable) |
| Compute | CPU-friendly; optional GPU accelerates embedding and generation |
| Integration | REST API with token-streaming (SSE) endpoints; built on NexusFlow |
| Deployment | Fully on-premise and offline-capable; assurable via NexusTrust for oversight |
| Clusters | Cross-cutting — Clusters 1, 2, 3, 4, 5, 6 |
How It Works
NexusScholar processes document collections through a six-stage pipeline designed for accuracy and traceability:
1. Ingest
PDFs and documents are parsed, cleaned, and split into semantically meaningful chunks with metadata preserved — page numbers, section headings, and document identifiers.
2. Embed & Index
Each chunk is transformed into a dense vector embedding and stored in a per-document cache. The vector index enables fast approximate nearest-neighbour search across the entire collection.
3. Retrieve
When a question is asked, the system retrieves the most semantically relevant passages from the vector index — casting a wide net to ensure no critical evidence is missed.
4. Rerank
A cross-encoder reranker re-scores the retrieved passages against the original question, pushing the most precisely relevant evidence to the top — significantly improving answer quality over embedding-only retrieval.
5. Generate
The top-ranked passages are provided as context to a locally served LLM, which generates a natural-language answer grounded entirely in the source material — with page-level citations for every claim.
6. Transform
The prompt-driven toolkit enables further operations on the answer or source text: summarise, expand, simplify, abstract, or rewrite — giving users flexible control over the output format and depth.
Heritage & Lineage
Developed by VTG as the document-intelligence layer of the Nexus suite, generalising retrieval-augmented generation over research and technical document collections. NexusScholar is built on the NexusFlow reasoning core and assured through NexusTrust.
Target Calls: Horizon-wide horizontal capability — applicable to Cluster 1, 2, 3, 4, 5 and 6 calls requiring evidence synthesis, document intelligence, or trustworthy AI-assisted knowledge work.
Try NexusScholar
NexusScholar is available as a hosted service for evaluation and demonstration purposes. Upload your document collections and start asking questions with grounded, cited answers.
Ready to discuss your project?
Get in touch to explore how Venaka can help translate your challenges into deployable AI solutions.
Book a Call →