Ai2 open-sources AstaBrief 8B for cited science reports
Ai2 has released AstaBrief 8B, an open-weights model that turns a research question into a cited report in about 51 seconds, along with the training data behind it.

The Allen Institute for AI has released AstaBrief 8B, an open-weights model that turns a research question and a pile of retrieved literature into a cited report in a single pass. Ai2 published the weights, the training data and an example workflow on 2 October, and the model is already handling the fast path inside Asta, the institute's agentic tool for scientific work.
A small model built for citations, not chatter
AstaBrief starts from Qwen3-8B and is trained purely for scientific report writing. Feed it a question plus the excerpts a retrieval step has already pulled from the literature, and it writes the finished report in one go rather than assembling it section by section. That single-pass design is the reason it is quick: across the full Asta pipeline, Fast mode averages 51.1 seconds per report against 178.5 seconds for the Claude-powered Thinking mode it sits beside, roughly 3.5 times faster. Ai2 also points out that open weights let institutions run the model on their own infrastructure, which matters when a research question touches sensitive or unpublished work.
How Ai2 trained it on real questions
The training pipeline was built from real queries rather than benchmark prompts. Ai2 filtered user logs down to 90,000 research-focused questions, stripped out bot traffic and personal information, and then generated full target reports for 47,000 of them using its ScholarQA pipeline backed by a mix of Claude 3.5 and 3.7 Sonnet, o3, o4-mini and GPT-4.1. For the preference stage it built roughly 6,000 comparison pairs and had GPT-4.1 and DeepSeek-R1 judge them, keeping only the pairs where both judges agreed. Ai2 says that agreement tracks human preference about 95 per cent of the time, and that insisting on it removed much of the noise that usually creeps into preference data.
The catch Ai2 puts in writing
Most of the training and evaluation was completed in 2025, so the proprietary models used as teachers and as comparison points reflect the frontier of that moment rather than today's. Ai2 states plainly that it has not rerun the full evaluation against current frontier models, and says the results should be read as evidence about its training and system-design choices rather than a claim about where an 8B model now sits. On the benchmarks it did run there is a real result: in a small human study, three scientific researchers ranked reports on preference, completeness, relevance, organisation and citation accuracy, and two of the three preferred AstaBrief on citation accuracy. Early usage is modest but encouraging, with 29.1 per cent of the 374 people who tried Fast mode returning on two or more days and users averaging 3.67 report threads each.
Our opinion
The interesting part of AstaBrief is not the benchmark table, it is the admission attached to it. Ai2 ships an 8B model that is roughly 3.5 times faster than the proprietary pipeline it replaces, then tells you the evaluation is a year old and the frontier has moved. That kind of note is rare and it makes the release more useful, not less, because the transferable work here is the data discipline: filtering 90,000 real queries, teaching the model to keep a claim inside what its source actually supports, and refusing preference pairs that two judges could not agree on. Citation-aware report writing is exactly the kind of task where a small, downloadable model beats renting a frontier one, especially for labs that cannot send unpublished results to somebody else's cloud. Ai2 has handed them the weights, the data and the recipe, which is the honest version of open.