News

Making drug pipeline data generally available

Today we’re launching a free, structured database of the global drug development pipeline (link to our platform here). Our goal was to consolidate information on thousands of pipeline drugs, targets, indications, and companies into an ergonomic interface designed for both humans and AI agents. 

You can use our datasets to help answer questions like:

  • Which drugs are being developed against a particular target?

  • Who are the main competitors for a new development program?

  • Which companies are active in a given disease?

  • What has previously been tried in an indication?

  • What are the upcoming catalysts for a therapeutic program?

  • Which mechanisms have advanced into the clinic, and which have failed?

Try it out by creating an account at platform.convoke.bio, and access the data via browser or our MCP server (which will be available on Claude, Phylo and Benchling's MCP marketplaces).

We see this release as a natural continuation of our work to build data and software infrastructure that improves the efficiency of the biopharmaceutical industry. A few months ago we launched the unmet needs index. We heard from many biotechs afterwards that they valued that someone had gone through the effort of cleaning up that information and making it easily accessible and digestible. A few have even used it to help them prioritize drug programs.

Both the pipeline and unmet need datasets are intended to help answer one of the most important questions in drug development: what potential medicines are worth working on? Pipeline project data helps answer this question in two main ways: competitive landscaping, and as a record of historical success and failure.

It’s still surprisingly difficult to identify who is working on or has historically worked on an analogous molecule. Even though most of the information required to question is (quasi-)public, it is fragmented across trial registries, company websites, scientific papers, conference abstracts, regulatory documents, investor presentations, and press releases. Pulling it together accurately and comprehensively is tedious and time-consuming — even with LLMs. But, it’s necessary; when working in crowded areas, proper knowledge and reactions to competitor movements is existential for new programs.

Of course, tools like ChatGPT and Claude have made this landscaping task much easier than it was a few years ago, but they still miss programs or pull from bad sources. And even if their responses are accurate, frontier LLMs are blunt instruments — repeatedly reconstructing basic facts from unstructured documents is unreliable and consumes excess tokens (money).

At Convoke, we build digital knowledge workers that help pharmaceutical and biotechnology companies make development decisions. When we found that we were repeatedly reconstructing therapeutic area pipeline information, we set out to make a simple queryable database that we could use to improve the performance of our agents.

We’ve seen that building this database and making it available to our agents can substantially improve performance of frontier models with web search alone on hard tasks, for instance ones that require agents to identify a comprehensive list of drug assets that fit specific criteria e.g., “Compile the complete list of active BCMA-targeted bispecific antibody programs in multiple myeloma”.

Adding the Convoke MCP tool to frontier models improves performance on complex tasks. Task difficulty was determined by the average percentage of relevant entities web-only agents were able to surface relative to a comprehensive ground truth (web recall). Evaluations were done with six total models from Anthropic, OpenAI, and Google, ranging from budget to frontier tiers. Error bars indicate ±1 SEM for each bucket.

Historically, assembling and maintaining a global drug pipeline dataset was expensive and wouldn’t be feasible for a startup. It required organizations of hundreds of people to search for new information, extract it manually, reconcile conflicting records and map it to a common taxonomy. The economics of that process produced an industry of high-priced data vendors whose products are largely inaccessible to smaller biotechnology companies, academic groups, patient organizations, and independent researchers. Now, technology can reproduce that work.

This an example of a broader transition in the life sciences: work that was previously labor-intensive, expensive, and restricted to the largest organizations is becoming programmable.

A great deal of knowledge work in biopharma consists of finding documents, extracting facts, reconciling terminology, and presenting the result in a form that supports a decision. Language models are beginning to automate each part of that process. As they do, capabilities that were once available only to the best-resourced pharmaceutical companies will become accessible to much smaller teams.

We hope that our database will serve as a community resource; a form of open shared memory that helps teams pick better problems to work on. 

Reality is messy, and no automated system—and no conventional database—is perfectly accurate. For instance, we know there are gaps in preclinical coverage. We will continue improving the data and the underlying extraction systems. We also hope the community will help us turn it into a better shared resource by leaving feedback. When you find an incorrect or incomplete record, please use the Share feedback button in the platform. Corrections will help us improve both the dataset and the systems that maintain it.

If you’re in SF, we are sponsoring a hackathon on August 13th where we can help you build agents with our MCP. 

If the data could be useful in your product—or you want to build agents that help your team accelerate drug development—get in touch at contact@convoke.bio.

Learn how we help teams unlock capacity

Stay updated

Copyright © 2026 Convoke Holdings, Inc.

All rights reserved.

Learn how we help teams unlock capacity

Stay updated

Copyright © 2026 Convoke Holdings, Inc.

All rights reserved.

Learn how we help teams unlock capacity

Stay updated

Copyright © 2026 Convoke Holdings, Inc.

All rights reserved.

Learn how we help teams unlock capacity

Stay updated

Copyright © 2026 Convoke Holdings, Inc.

All rights reserved.