
A new capability in Emet that takes a research goal and returns novel, mechanism-bearing therapeutic hypotheses—with provenance attached. We tested four of its proposals for lung squamous cell carcinoma in human cells; each happened to involve modulation of two targets. Two combinations produced effects in vitro exceeding their single agent activities. Neither combination had been tested together in this disease before.
Quick links
Request an Emet briefing for your discovery team
Sign up for the webinar to see a live demo
Today, we're introducing the Hypothesis Generation Agent
The Hypothesis Generation Agent is a new capability inside Emet. You give it a research goal in natural language — find novel therapeutic targets in this disease — and it returns ranked, mechanistically explicit hypotheses, each one carrying the reasoning that produced it and the evidence that supports it.
Not a literature summary. Not a ranked list of genes. A hypothesis with a mechanism: these two targets, in this genotype, for this reason, and here is what should happen if the reasoning is right.
In the paper we published, we call it the Agentic Research Director, because that is what it does — it works the way a human research director does, running many lines of inquiry in parallel, giving each one a team of tools, learning from what comes back, and deciding what to pursue next.
This post covers both halves: the science we ran to test whether it works and why Emet is the environment in which a capability like this can exist at all.
Why does this have to sit on top of a fused stack?
The scientific stack is fragmenting faster than any scientist can keep up with. Thousands of data sources. 700+ models. Endless software, agents, omics platforms, and lab automation. It only gets worse from here.
But the value was never in the pieces. It's in bringing them together. No single model or dataset produces a novel therapeutic hypothesis. Fusing them does.
Emet is a domain-specific AI for drug discovery — a biology-grounded platform that fuses the fragmented scientific stack into a single system that thinks like a scientist. It spans 200 R&D workflows and brings together data, every frontier, and specialized models, software, and lab-in-the-loop in one place. Emet acts as the conductor of your science: it connects to your internal data, models, software licenses, and cloud, and is tuned by our scientists to your therapeutic areas, your compliance requirements, and the way your teams actually work.
The Hypothesis Generation Agent is what that fusion makes possible. It is not a model you could license separately and bolt onto an existing pipeline. It only works because it can reason across multi-omic, functional, pharmacological, clinical, and literature evidence inside a single environment — which is exactly what produced the two results below.
How the Hypothesis Generation Agent works
Emet is built in four layers, and the agent uses all of them.
The data layer is a biomedical knowledge graph of more than 1.5 billion triples, with relationships derived from over 6 million papers, canonicalized to a common identity system and connected to more than 100 specialized biological databases — literature (PubMed, Europe PMC, OpenAlex), cancer genomics (cBioPortal, TCGA, DepMap), expression (GTEx, single-cell atlases), variants (ClinVar, gnomAD), pharmacology (ChEMBL, Open Targets), structure (AlphaFold, PDB), pathways (Reactome, STRING), trials, and patents.
The model layer exposes frontier and specialized models behind one interface, including a custom graph neural network for link prediction. The orchestration layer composes them into workflows. The validation layer cross-references every surviving claim against the literature and experimental evidence base, keeps the provenance, and deprioritizes hypotheses that are inadequately supported or not actually novel.
Three design choices in the agent itself matter for the result.
It searches a graph of hypotheses, not a single chain of reasoning. Each node is one hypothesis pursued under one strategy — synthetic-lethal dependency, protein-interaction convergence, expression-based vulnerability, precedent from validated combinations in other indications. Nodes generate, refine against supporting and disconfirming evidence, and score for mechanistic plausibility. Weak nodes get pruned rather than carried forward. Independent lines of reasoning branch, run in parallel, and recombine where they converge on the same mechanistic logic.
It uses models as scoped instruments, not global ranking engines. This is the departure from conventional knowledge-graph pipelines, and it is what made these results reachable. Standard practice runs a predictive model once over the entire space — every gene–disease pair, every drug pair — producing one context-free ranking fixed before any biological question is asked. A candidate's visibility then depends on its position in a universal ordering. The agent instead samples the region of the graph its current strategy makes relevant, specifies exactly what it wants predicted over that region, and interprets the answer relative to the line of reasoning that requested it. The same model gives different, locally relevant answers depending on which branch is asking.
Its control loop is deterministic; the models are confined to judgment. Which node gets worked on next, how child scores aggregate to parents, what gets committed to memory, and when to stop are all code, not model discretion. A separate critic reviews the evidence — deliberately not the same process that produced and self-scored it. Every round's strategies, scores, keep-or-drop decisions and already-known art are written to persistent memory inside the loop, so later rounds extend what worked, skip what has already been exhausted, and leave an auditable trace of what was discarded and why.
For a regulated R&D organization, that last point is not a technical footnote. It's the difference between a hypothesis you can defend in a portfolio review and one you can't.
What it found, and what the lab said
A complex disease: the right first test. Lung squamous cell carcinoma (LUSC) is the canonical case explored here. LUSC comprises 25–30% of non-small cell lung cancer with five-year survival under 10% once metastatic, and is one of the most mutationally complex solid tumours according to TCGA. There is near-universal TP53 loss, but also recurrent amplification of SOX2, PIK3CA and FGFR1 and many other driving mutations. There is no dominant druggable driver given the multiple points of dysregulation across key cellular processes. Hence the long record of failure - FGFR inhibitors like AZD4547 and rogaratinib; anti-angiogenics closed off in squamous histology after fatal pulmonary haemorrhage; PI3K monotherapy, despite frequent PIK3CA amplification, to name a few. Designing a strategy here means weighing genetics, checkpoint biology, understanding all points of dysregulation in these cells, redundancies and potentially smartly leveraging the high mutational burden of cells for a therapeutic advantage - the conjunction manual search handles worst.
We asked one question of Emet and the Agentic Research Director: identify novel therapeutic targets in Lung squamous cell carcinoma (LUSC). No initial target list, no drug class, no instruction to look for specific therapeutic strategies.
Emet returned dozens of evidence-validated hypotheses; the top 20 were vetted internally, the top four sent to the wet-lab. Unprompted, every one of the top four was a two-target combination. For a disease with multiple drivers, that was an interesting emergent behavior of the system’s reasoning. Two produced effects exceeding single-agent treatment, neither previously tested together in LUSC to our knowledge.
CDC7 + PKMYT1. CDC7 licenses and fires replication origins via MCM2–7; PKMYT1 phosphorylates CDK1 at Thr14, holding CDK1–cyclin B inactive through S phase and G2. Inhibit CDC7 and replication is left incomplete; inhibit PKMYT1 and the brake that would arrest those cells before division is gone, so they enter mitosis carrying under-replicated chromosomes. The two non-overlapping checkpoints converge on mitotic catastrophe. LUSC is unusually permissive because near-universal TP53 loss has already dismantled the G1 checkpoint.
Conventional methods would have struggled. Never tested together in any LUSC line, the pair sits beyond supervised synergy prediction, and no screen would prioritise it: CDC7 and PKMYT1 neither interact physically nor share a cell-cycle phase, and their convergence is visible only through CDK1 regulation, and only after TP53 loss removes G1.
We tested TAK-931 against RP-6306 in 7×7 dose matrices in NCI-H520 and SK-MES-1, and got statistically significant synergistic growth inhibition in both lines, strengthening over time: HSA scores rose from 19.24 to 24.19 between day 5 and day 7 in NCI-H520, and 11.07 to 13.73 in SK-MES-1. That is what an effect compounding across cell cycles should look like — and it came unoptimised, with no dosing sequence or biomarker selection explored. Optimization along those factors will likely further amplify the effect.
USP13 + PI3Kα. USP13 is among the most frequently amplified genes in LUSC, on the 3q26 amplicon alongside PIK3CA and SOX2. As a deubiquitinase it stabilises c-MYC and the squamous lineage programme, and in other epithelial cancers MCL-1; PI3K/AKT/mTOR signalling independently drives de novo production of the same effectors. One arm acts on protein stability, the other on signal transduction, so co-inhibition should deplete the standing reserve of pro-survival protein while blocking replenishment. The aim in this case was not supra-additive killing but reopening a therapeutic window for a class that has repeatedly failed as LUSC monotherapy, something a synergy score cannot detect and a synergy-ranked screen would discard.
Both predictions held, most clearly in NCI-H520. Alpelisib with spautin-1 gave additive inhibition at 48 h there and a much weaker interaction in SK-MES-1. In the NCI-H520 matrix, 30 µM alpelisib alone gave ~51% inhibition, while comparable inhibition came from 15 µM alpelisib plus 20 µM spautin-1. So, half the PI3K inhibitor dose produced the same effect with a USP13 inhibitor. The predicted mechanism was then confirmed: immunoblotting at 24, 48 and 72 h showed MCL-1 falling to 0.31 and 0.33 of control by 48 h and to 0.25 and 0.01 by 72 h, SOX-2 to 0.47 then 0.32 while alpelisib alone left it unchanged, and c-Myc declining alongside. The effect tracked the USP13 arm, not the PI3K arm, exactly as predicted. The proxies for a lowered apoptotic threshold and a disrupted squamous lineage switch were written down first, then observed.
One question, and the many others shaped like it. This is one disease and two cell lines; it points to possibility, raises interest. , Clinical promise is not in play yet. What it shows is the value of an integrated environment with reasoning power enough to find signal across many datasets when pointed at one specialised question, and to return candidates carrying their mechanism rather than a score alone. The same substrate extends naturally: drug repurposing is this query in reverse, asking which diseases an agent’s mechanism should address; indication expansion is the narrower version, attractive because it inherits a known clinical profile and shortens the path to trial. Both benefit from what mattered most here: an explicit mechanism specifies which patients should respond, and how a trial should be stratified.
Based on “Novel Target Combinations in Lung Squamous Cell Carcinoma proposed by the Emet AI Research Environment and supported by discovery stage experimentation” — Soman, Newington et al. In vitro work by Pharmaron.
Getting access
The Hypothesis Generation Agent will be available to Emet customers in the next couple of weeks, tuned by our scientists to your therapeutic areas and connected to your internal data, models, and software.
In vitro experiments were performed under contract by Pharmaron (Beijing, China) to protocols designed by the authors, who specified the treatment groups and analysis plan and interpreted the resulting data.
Want to explore what’s possible?
Try EMET

