Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
AstaBrief open-source model speeds scientific report generation - OpenSmartRoute
AstaBrief open-source model speeds scientific report generation
AstaBrief 8B, a small open model for scientific report writing, is now available. It generates cited reports faster and with comparable quality to proprietary models.
Key points
AstaBrief 8B reduces report generation time to 51.1 seconds from 178.5 seconds.
The model is based on Qwen3-8B and trained with real scientific queries.
Training used over 90,000 research-focused queries and 47,000 quality-filtered examples.
Open-sourcing allows institutions to run AstaBrief locally for sensitive data.
Why it matters: Faster, open models improve scientific workflows and reduce costs while maintaining report quality and evidence grounding.
By OpenSmartRoute editorial · written through the router by llm-small
From Hugging Face blog - “Open-sourcing AstaBrief, the fast report-generation model in Asta”
Introduction - Announcing AstaBrief as an open-source, fast report-generation model
AstaBrief is a new open-source model designed for scientific report writing. It can generate detailed, cited reports quickly. The model is small enough to run on local computers. It aims to match the quality of larger, proprietary models while being faster and more accessible. Researchers can now download and use AstaBrief themselves. This makes scientific report generation more open and flexible. The model is available in Asta’s platform as a fast mode option. It is also open-sourced, along with the training data. This allows others to study, reproduce, and improve on the approach.
Background - Scientific report needs and current AI limitations
Scientific work often requires detailed reports that cite evidence accurately. Researchers want AI tools that help find literature, synthesize findings, and answer complex questions. However, current language models face challenges in meeting these needs. They can generate text quickly but often lack proper grounding in evidence. Sometimes, they make claims that are not supported by the sources they cite. This can lead to inaccuracies and reduce trust in AI-generated reports. Researchers also need to verify and trace the evidence behind each statement. Existing models often produce reports in sections, which can be slow and inefficient. They may also struggle to incorporate large amounts of context or constraints from scientists. These limitations highlight the need for faster, more reliable scientific report tools.
Model architecture - Based on Qwen3-8B, designed for scientific synthesis
AstaBrief is built on a model called Qwen3-8B. This is a smaller language model with 8 billion parameters. It was chosen because of its balance between size and performance. The model was adapted specifically for scientific synthesis tasks. It is designed to turn a research question and retrieved literature into a complete, cited report. Unlike larger models, AstaBrief can generate reports in a single pass. This means it writes the entire report at once, rather than section by section. The architecture focuses on understanding the question, retrieving relevant evidence, and synthesizing it into a structured report. The goal was to create a model that is both fast and capable of producing high-quality, grounded scientific text.
Google launched EmbeddingGemma 2, a compact open model that handles text, code, images, video, and audio. It uses a single shared vector space to enable unified search across all media types.
Training data - Real user queries, literature retrieval, and filtering methods
Training AstaBrief involved collecting real research queries submitted by scientists. These queries often include substantial context and multiple constraints. Researchers analyzed hundreds of thousands of queries to understand what scientists ask AI tools. They then filtered these queries to remove irrelevant or low-quality requests. Filtering included dropping non-English, personal, or non-scientific prompts. This process left about 90,000 research-focused queries for training.
The training data also included literature retrieval results. The system retrieved relevant scientific papers and organized the information into sections. For supervised fine-tuning, the team generated full reports from the filtered queries. They used a pipeline that retrieved literature, structured the material, and synthesized evidence into a report. The training examples came from a mix of proprietary systems and multiple language models. They also created pairs of reports with different qualities for preference training. These pairs helped the model learn what makes a good, well-supported report.
Training process - Supervised fine-tuning and preference optimization techniques
The training process combined two main methods: supervised fine-tuning (SFT) and direct preference optimization (DPO). SFT involved teaching the model to generate full reports from real queries and literature. The team used a set of high-quality examples created by existing systems. These examples helped the model learn how to synthesize evidence into a coherent report.
DPO involved comparing pairs of reports and teaching the model which one was better. They generated pairs from different sources and used judge models to select the preferred report. The judges were large language models trained to align with human preferences. They only kept pairs where both judges agreed, ensuring higher quality training data. This method helped the model learn to produce reports that are both relevant and well-supported. The combination of SFT and DPO aimed to improve the model’s ability to generate accurate, grounded scientific reports efficiently.
Evaluation - Benchmarks, metrics, and comparison with proprietary models
The model was tested using several benchmarks and metrics. The main evaluation was on a set of 200 research questions called SQABench-CS2. This benchmark measures how well the report covers the necessary content, relevance, and citation accuracy. It includes four key metrics: coverage of content, relevance of answers, accuracy of citations, and whether citations support the claims.
Secondary evaluations used additional benchmarks, such as DeepScholarBench, which contains 63 recent research questions. They also compared AstaBrief’s reports with those generated by proprietary models. Judges, including large language models and humans, assessed the quality of the reports. The goal was to see if AstaBrief could match or surpass the quality of larger, closed models while being faster. The results showed that AstaBrief’s report generation time was significantly shorter, averaging about 51 seconds per report compared to 179 seconds for the proprietary pipeline. This made AstaBrief about 3.5 times faster overall.
Grounding in evidence - Ensuring citations support claims accurately
Grounding reports in evidence is crucial for scientific trustworthiness. AstaBrief was designed to ensure citations support the claims made in the report. The model was trained to attach citations that are relevant and directly support the statements. However, citations alone do not guarantee accuracy. A model can cite the right study but still overstate or generalize findings. For example, it might turn a specific result into a broad conclusion or change the tense to make a statement sound more universal.
To address this, the development focused on measuring how well citations support the claims. They used metrics to evaluate relevance, coverage, and citation grounding. These metrics help identify whether the report’s statements are backed by appropriate evidence. The team also tested for issues like unsupported claims or overgeneralizations. Improving citation grounding helps ensure that reports are both accurate and trustworthy for scientific use.
Implications - Benefits for scientific research and local deployment
Open-sourcing AstaBrief offers many benefits for scientific research. Researchers can run the model on their own infrastructure, which is important for sensitive or unpublished work. Local deployment allows scientists to generate reports without sharing data externally. This increases privacy and control over research data.
The model’s speed also makes it useful for quick, preliminary reports. Scientists can generate initial drafts and then refine them. The open approach encourages collaboration and customization. Institutions can adapt AstaBrief to specific fields or workflows. The release of training data and example workflows provides a starting point for local report generation. This helps make scientific synthesis more accessible and efficient across different research communities.
How it compares - what existed before, what this changes and what stays the same
Before AstaBrief, most scientific report generation relied on proprietary models or manual writing. These models often used large, closed-source systems that were expensive to run. They also required multiple steps, such as summarization, clustering, and section-by-section writing. This process was slow and less transparent.
AstaBrief introduces a smaller, open-weight model trained specifically for scientific report generation. It can produce full reports in one pass, which reduces the time needed. The model is designed to be fast and accessible for local deployment. It is trained on real user queries and literature excerpts, making it more aligned with actual scientific needs.
What stays the same is the importance of citations and evidence grounding. The goal remains to produce reports that are accurate, relevant, and well-structured. The underlying principles of using retrieval-augmented generation (RAG) and fine-tuning still apply. The focus on grounding reports in evidence and providing verifiable outputs continues to be central.
AstaBrief differs mainly in its size, speed, and openness. It is smaller than many proprietary models but still aims to match their report quality. Its open-source nature allows institutions to run it on their own hardware, which was not possible with many previous solutions. This change makes scientific report generation more accessible, transparent, and customizable.
Questions this leaves open - what the source does not say and how a reader can check it
The source does not specify the exact size of the training data or the diversity of scientific fields covered. It mentions tens of thousands of real user queries, but not how representative they are across disciplines. Researchers might want to know if the model performs well in their specific area.
It also does not detail the evaluation metrics used to measure report quality. While it mentions relevance and citation grounding, the exact benchmarks or comparison results are not provided. This makes it harder to assess how AstaBrief stacks up against other models in different tasks.
The source does not clarify how well the model handles conflicting evidence or ambiguous literature. Scientific research often involves uncertain or contradictory findings. It is unclear how AstaBrief manages these complexities or whether it can flag uncertain areas.
Another open question is about the licensing and usage restrictions. The model is open-source, but the license type and any limitations are not specified. Users may want to know if they can modify, commercialize, or redistribute the model freely.
The source also does not describe how the model can be further improved or customized. For example, can users fine-tune it with their own datasets? How easy is it to adapt the model to new scientific domains? These are important considerations for practical deployment.
To check these aspects, users can run their own evaluations. They can test the model on their specific datasets or literature. Comparing its outputs with human-written reports can reveal strengths and weaknesses. Reviewing the training data and evaluation code, if available, can also provide insights into the model’s capabilities and limitations.
Finally, engaging with the open-source community or the developers can help clarify unresolved questions. Feedback and shared experiences can guide future improvements and best practices for using AstaBrief effectively.
What to do - How researchers and institutions can adopt AstaBrief
Researchers interested in using AstaBrief can download the open-source model and training data. They can run the model on their own hardware, provided they meet the system requirements. The example workflows included with the release can be adapted to different types of literature, such as PDFs or databases.
Institutions can integrate AstaBrief into their research tools or pipelines. They can also fine-tune the model further with their own data to improve accuracy in specific fields. The open-source nature allows for transparency and continuous improvement. Researchers and organizations can contribute to the development by testing, providing feedback, or sharing new training data. This approach helps build better tools for scientific report generation and supports open science.