Insights

Benchmarking biopharma knowledge work

At Convoke, we help our biopharma customers transform their business with AI. To do that successfully in a production system, we need to deeply understand the performance of different agents and models for all of the work that might happen within a biopharma company. It’s a challenge to select the right model for each task as the capability frontier is jagged; even leading LLMs can produce excellent work on one task and useless slop on another.

Today we’re releasing the first set of results from our biopharma knowledge work benchmark. Much of the effort in benchmarking capabilities of models in the life sciences has gone towards research and development (primarily drug discovery) use cases, but relatively little towards benchmarking capabilities in activities that support the commercial, development, or operational functions of a biopharmaceutical company (the customers served by Convoke). This is the gap we hope to fill with this benchmark.

These initial results evaluate model performance across 6 diverse biopharma knowledge work tasks, but we have identified 500+ work products that are common in biomedical enterprises and we intend to expand our benchmark to cover substantially all of these in time.

Goals and design philosophy

There are hundreds if not thousands of potential use cases for LLMs in a biopharma enterprise that are common enough to justify automation, and without a quantitative benchmark it’s difficult to convincingly answer these common questions:

  • What use cases should we prioritize?

  • Where is frontier intelligence necessary? Where are cheaper models sufficient?

  • Where are off the shelf models good enough, and where do they need additional harnesses and/or custom builds?

  • How are capabilities changing over time? Can we wait for models to catch up?

The main design goals of our benchmark were:

  • A scalable and extensible methodology that works across a diversity of tasks

  • Sufficient sensitivity to discriminate where models are performing well or not for a broad and diverse range of tasks

  • A scoring system that is generalizable across many tasks that allows us to easily add new tasks and grading schemas, or extend existing grading schemas

Methodology

We adopted a penalty-based methodology where each task has its own checklist-like grading schema created by life science domain experts. 

We followed this process to generate the grading schemes:

  1. We selected a small number of work products that are common among our customers, and diverse in format and analysis

  2. Domain experts enumerated an extensive list of errors for each task type based on domain expertise and observed issues seen in practice with LLM generations

  3. Then, these experts determined what the contents and properties of a high-quality output would be for each work product

  4. Experts synthesized the above information into a detailed and specific error-based grading schema that could be evaluated by an LLM grader agent

The work products we chose to evaluate for this first version of the benchmark are:

  • Target safety assessment: A risk assessment summarizing safety evidence, potential toxicities, and liabilities for a target

  • Tumor landscape: A report mapping a cancer indication’s patient segments, treatment pathways, competing drugs, and unmet needs

  • Shelved assets: A shortlist of discontinued or paused drug programs, explaining why they stopped and their potential for revival and/or inlicensing

  • Catalyst calendar: A dated tracker of upcoming trial readouts, regulatory decisions, and other events that could shift the competitive landscape

  • Clinical benchmark: A comparison table of relevant trials and therapies, establishing efficacy and safety thresholds a new drug should meet

  • Commercial forecast: A spreadsheet model projecting patient uptake, pricing, market share, and revenue under different assumptions

These provide a mix of output formats (spreadsheets, narrative reports, powerpoint decks) and types of analysis (compiling comprehensive datasets, synthesizing literature). Most of these tasks require extensive research and report generation and directly test agentic capabilities and long-running tasks.

For each work product we prepared a grading schema with penalties for potential errors or omissions. As an example of some of these statements an excerpt from the target safety assessment grading scheme is included below:

  • The report does not contain a summary of the genetics of the target and its protein sequence and structure (🔴Major error)

  • The report does not evaluate possible differences in biological function between humans and common model species (🔴Major error)

  • The report does not enumerate potential vitro or pharmacologically relevant animal models available for safety testing (🟠Moderate error)

  • The report does not evaluate similarities and differences in gene and protein sequence across humans and model species (🟠Moderate error)

  • The report does not evaluate differences in the binding epitope, active sites, allosteric sites, or other key regions across species if relevant (🟡Minor error)

Example statements from the Target Safety Assessments rubric related to biological sequence and function

Our grading methodology is somewhat analogous to an end-to-end testing framework in software engineering in that we do not tightly specify how a model achieves some task, as long as it passes prespecified criteria. In testing, we gained confidence based on expert review of grader outputs that this error-based approach was able to score outputs in a manner that revealed differences in performance across tasks and across models within each task.

We used OpenCode harness with Parallel MCP to evaluate the models. Models were run at their respective max effort levels. Agents were instantiated in a sandbox, where they were given basic instructions on how to navigate the file system, where to output the final artifact, the current date, and a biopharma prompt. We kept the prompts sparse, because we wanted to evaluate the models on their innate ability to resolve biopharma tasks.

To evaluate the outputs, we gave grading agents access to the artifact, a description of the prompt, and Convoke-proprietary tools with which to perform research. The agents were able to judge artifacts on their internal inconsistencies. They were also able to grade artifacts based on external research. Each criterion grading was rooted in a citation. 

Results

Performance

Overall, the frontier OpenAI models lead across tasks (GPT-6 Astra, GPT-6- and GPT-5.6 Sol, GPT-5.6 Terra, and GPT-6 Luna) with open weight and Gemini models performing much more poorly. Opus 5 outperformed other frontier models on one task, but could not compete with latest era OpenAI models across all tasks. We could not benchmark Opus 5.5 or Fable due to restrictions on biological content.

We also see large differences in performance across task types. Frontier agents do very well at tasks that require them to compile structured Excel files (shelved assets, catalyst calendars). These tasks are somewhat brute-forceable through extensive search and dataset composition, and require little narrative synthesis. Performance is notably worse at tasks that require narrative reports, persuasion, and integrating evidence (target safety, tumor landscape reports). While commercial forecast models are an Excel output, they require judgement in design, assumptions, and are sensitive to errors. 

In general, tasks that we could classify as ‘integration’ or ‘synthesis’ tasks are harder for models than tasks that can be completed by iteratively compiling and appending information to a report or database. This is perhaps not surprising given the autoregressive nature of how LLMs generate outputs (in a stream); this weakness can generally be addressed with harness design.

Cost

Frontier lab models were substantially more expensive than all other tested models (>3x) per work product generation. In general, OpenAI and Anthropic models were the most expensive. When visualized as a pareto-style chart, GPT-5.6 Terra and GPT-6 Luna offer an attractive tradeoff of cost versus performance.

Conclusion and next steps

We expect life science knowledge work will grow in importance for future model and agent releases.  We are continuing to expand the tasks covered by the benchmark, as well as improving the rubrics and grader, and expect to cover hundreds of task types in the coming months. Based on the preliminary results, we recommend frontier OpenAI models for when performance is critical and GPT-6 Luna when aiming to contain costs.

Our future goals for the biopharma knowledge work benchmark include:

  1. Further improve grader agent

    1. Create expert-created ‘golden outputs’ for each task to compare against LLM generated outputs

    2. Validate the grader agent by evaluating concordance with expert grading

  2. Extending the existing rubrics

    1. Refine grading schema following thorough review of the batch of generated work products and grader feedback

  3. Add additional tasks

    1. Work through our list of biopharma 500+ work products to add additional grading schema

  4. Evaluate harness performance, including the Convoke Coworker harness

We are using these and our internal results to support our customers in implementations of agents in their business. If you are interested in seeing more details or understanding the performance of another model, please reach out to contact@convoke.bio.

Learn how we help teams unlock capacity

Stay updated

Copyright © 2026 Convoke Holdings, Inc.

All rights reserved.

Learn how we help teams unlock capacity

Stay updated

Copyright © 2026 Convoke Holdings, Inc.

All rights reserved.

Learn how we help teams unlock capacity

Stay updated

Copyright © 2026 Convoke Holdings, Inc.

All rights reserved.

Learn how we help teams unlock capacity

Stay updated

Copyright © 2026 Convoke Holdings, Inc.

All rights reserved.