Insights
Benchmarking biopharma knowledge work

At Convoke, we help our biopharma customers transform their business with AI. To do that successfully in a production system, we need to deeply understand the performance of different agents and models for all of the work that might happen within a biopharma company. It’s a challenge to select the right model for each task as the capability frontier is jagged; even leading LLMs can produce excellent work on one task and useless slop on another.
Today we’re releasing the first set of results from our biopharma knowledge work benchmark. Much of the effort in benchmarking capabilities of models in the life sciences has gone towards research and development (primarily drug discovery) use cases, but relatively little towards benchmarking capabilities in activities that support the commercial, development, or operational functions of a biopharmaceutical company (the customers served by Convoke). This is the gap we hope to fill with this benchmark.
These initial results evaluate model performance across 6 diverse biopharma knowledge work tasks, but we have identified 500+ work products that are common in biomedical enterprises and we intend to expand our benchmark to cover substantially all of these in time.
Goals and design philosophy
There are hundreds if not thousands of potential use cases for LLMs in a biopharma enterprise that are common enough to justify automation, and without a quantitative benchmark it’s difficult to convincingly answer these common questions:
What use cases should we prioritize?
Where is frontier intelligence necessary? Where are cheaper models sufficient?
Where are off the shelf models good enough, and where do they need additional harnesses and/or custom builds?
How are capabilities changing over time? Can we wait for models to catch up?
The main design goals of our benchmark were:
A scalable and extensible methodology that works across a diversity of tasks
Sufficient sensitivity to discriminate where models are performing well or not for a broad and diverse range of tasks
A scoring system that is generalizable across many tasks that allows us to easily add new tasks and grading schemas, or extend existing grading schemas
Methodology
We adopted a penalty-based methodology where each task has its own checklist-like grading schema created by life science domain experts.
We followed this process to generate the grading schemes:
We selected a small number of work products that are common among our customers, and diverse in format and analysis
Domain experts enumerated an extensive list of errors for each task type based on domain expertise and observed issues seen in practice with LLM generations
Then, these experts determined what the contents and properties of a high-quality output would be for each work product
Experts synthesized the above information into a detailed and specific error-based grading schema that could be evaluated by an LLM grader agent
The work products we chose to evaluate for this first version of the benchmark are:
Target safety assessment: A risk assessment summarizing safety evidence, potential toxicities, and liabilities for a target
Tumor landscape: A report mapping a cancer indication’s patient segments, treatment pathways, competing drugs, and unmet needs
Shelved assets: A shortlist of discontinued or paused drug programs, explaining why they stopped and their potential for revival and/or inlicensing
Catalyst calendar: A dated tracker of upcoming trial readouts, regulatory decisions, and other events that could shift the competitive landscape
Clinical benchmark: A comparison table of relevant trials and therapies, establishing efficacy and safety thresholds a new drug should meet
Commercial forecast: A spreadsheet model projecting patient uptake, pricing, market share, and revenue under different assumptions
These provide a mix of output formats (spreadsheets, narrative reports, powerpoint decks) and types of analysis (compiling comprehensive datasets, synthesizing literature). Most of these tasks require extensive research and report generation and directly test agentic capabilities and long-running tasks.
For each work product we prepared a grading schema with penalties for potential errors or omissions. As an example of some of these statements an excerpt from the target safety assessment grading scheme is included below:
|
Example statements from the Target Safety Assessments rubric related to biological sequence and function
Our grading methodology is somewhat analogous to an end-to-end testing framework in software engineering in that we do not tightly specify how a model achieves some task, as long as it passes prespecified criteria. In testing, we gained confidence based on expert review of grader outputs that this error-based approach was able to score outputs in a manner that revealed differences in performance across tasks and across models within each task.
We used OpenCode harness with Parallel MCP to evaluate the models. Models were run at their respective max effort levels. Agents were instantiated in a sandbox, where they were given basic instructions on how to navigate the file system, where to output the final artifact, the current date, and a biopharma prompt. We kept the prompts sparse, because we wanted to evaluate the models on their innate ability to resolve biopharma tasks.
To evaluate the outputs, we gave grading agents access to the artifact, a description of the prompt, and Convoke-proprietary tools with which to perform research. The agents were able to judge artifacts on their internal inconsistencies. They were also able to grade artifacts based on external research. Each criterion grading was rooted in a citation.
Results
Performance
Overall, the frontier OpenAI models lead across tasks (GPT-6 Astra, GPT-6- and GPT-5.6 Sol, GPT-5.6 Terra, and GPT-6 Luna) with open weight and Gemini models performing much more poorly. Opus 5 outperformed other frontier models on one task, but could not compete with latest era OpenAI models across all tasks. We could not benchmark Opus 5.5 or Fable due to restrictions on biological content.
We also see large differences in performance across task types. Frontier agents do very well at tasks that require them to compile structured Excel files (shelved assets, catalyst calendars). These tasks are somewhat brute-forceable through extensive search and dataset composition, and require little narrative synthesis. Performance is notably worse at tasks that require narrative reports, persuasion, and integrating evidence (target safety, tumor landscape reports). While commercial forecast models are an Excel output, they require judgement in design, assumptions, and are sensitive to errors.
In general, tasks that we could classify as ‘integration’ or ‘synthesis’ tasks are harder for models than tasks that can be completed by iteratively compiling and appending information to a report or database. This is perhaps not surprising given the autoregressive nature of how LLMs generate outputs (in a stream); this weakness can generally be addressed with harness design.
Cost
Frontier lab models were substantially more expensive than all other tested models (>3x) per work product generation. In general, OpenAI and Anthropic models were the most expensive. When visualized as a pareto-style chart, GPT-5.6 Terra and GPT-6 Luna offer an attractive tradeoff of cost versus performance.
Conclusion and next steps
We expect life science knowledge work will grow in importance for future model and agent releases. We are continuing to expand the tasks covered by the benchmark, as well as improving the rubrics and grader, and expect to cover hundreds of task types in the coming months. Based on the preliminary results, we recommend frontier OpenAI models for when performance is critical and GPT-6 Luna when aiming to contain costs.
Our future goals for the biopharma knowledge work benchmark include:
Further improve grader agent
Create expert-created ‘golden outputs’ for each task to compare against LLM generated outputs
Validate the grader agent by evaluating concordance with expert grading
Extending the existing rubrics
Refine grading schema following thorough review of the batch of generated work products and grader feedback
Add additional tasks
Work through our list of biopharma 500+ work products to add additional grading schema
Evaluate harness performance, including the Convoke Coworker harness
We are using these and our internal results to support our customers in implementations of agents in their business. If you are interested in seeing more details or understanding the performance of another model, please reach out to contact@convoke.bio.
