AstaBrief writes a cited report in 51.1 seconds
Ai2 has open-sourced AstaBrief 8B, which turns a research question and literature excerpts into a cited report in about 51 seconds.
Published · on allenai.org · 2 min read

Ai2 has open-sourced AstaBrief 8B, a model that takes a research question and retrieved literature excerpts and writes a cited report from them. The weights, the training data and an example workflow are published, and the model runs inside Asta, Ai2's agentic platform for scientific work, as Fast mode alongside a Claude-powered Thinking mode. Ai2 describes the release on allenai.org, in a post dated 2 October 2026.
51.1 seconds against 178.5 seconds
Across the full Asta pipeline, Fast mode averages 51.1 seconds per report, against 178.5 seconds for Thinking mode, about 3.5× faster. Ai2 redesigned the pipeline so the model writes the full report in one pass rather than section by section, bypassing the snippet summarization and clustering stages the Claude-based mode uses, and reports that it found this possible without sacrificing performance.
Training started from Qwen3-8B. Ai2 chose supervised fine-tuning and direct preference optimization over reinforcement learning, which it describes as unstable and expensive, and put most of the effort into data quality: 90K filtered research queries, 47K usable training examples, and about 6K preference pairs where two judge models had to agree on the winner. The main evaluation target was SQABench-CS2, 200 user-written computer science questions, tracking rubric score, answer precision, citation precision and citation recall. Ai2 also ran DeepScholarBench, a 63-query long-form synthesis benchmark, an LLM-judged pairwise comparison against the Claude-powered pipeline, and a small human study.
What open weights change
Open weights let institutions run AstaBrief on their own infrastructure, which Ai2 says is necessary when research questions reveal sensitive or unpublished work. The released example workflow can be adapted to generate reports from a researcher's own PDFs, giving a starting point for local report generation.
Ai2 notes that most of the training and evaluation was completed in 2025, so the proprietary models used to generate training data and as comparison points reflect the frontier at the time, and the full evaluation has not been rerun against the current frontier models. The published metrics cover relevance, coverage and citation grounding; Ai2 writes that a richer evaluation of scientific report writers should also test whether they preserve the scope and strength of the claims in their sources.
Source: allenai.org — BARGO’s commentary on the linked source.
Event date: