Skip to main content
Factlen ExplainerBiomedical AIEvidence PackJun 29, 2026, 11:40 PM· 5 min read· in data analysis

AI Model 'MAMMAL' Outperforms AlphaFold 3, Signaling a New Era for Multimodal Drug Discovery

IBM Research has released MAMMAL, a multimodal foundation model trained on 2 billion biological samples that achieves state-of-the-art results across nine drug discovery benchmarks. By integrating proteins, small molecules, and gene expression data, the open-source model outperformed AlphaFold 3 in specific antibody-binding tests, offering a new unified tool for biomedical research.

By Nicolas Laurent

Computational Biologists 35%Open-Source Advocates 25%Structural Biologists 20%Clinical Translators 20%
Computational Biologists
Argue that unified, multi-modal architectures are the future of drug discovery, moving beyond single-task models.
Open-Source Advocates
Emphasize the importance of releasing model weights and code publicly to democratize biotech research.
Structural Biologists
Maintain that while multimodal models excel at classification, dedicated structural models like AlphaFold remain essential for understanding physical mechanisms.
Clinical Translators
Focus on the reality that all computational predictions must ultimately survive rigorous wet-lab validation and human trials.

Key points

  • IBM Research introduced MAMMAL, a multimodal AI model for drug discovery.
  • The model integrates proteins, antibodies, small molecules, and gene expression data.
  • MAMMAL achieved state-of-the-art results on 9 out of 11 drug discovery benchmarks.
  • It outperformed AlphaFold 3 in specific antibody-antigen binding classification tasks.
  • The model weights and codebase are open-source, lowering the barrier for biotech startups.

The introduction of AlphaFold revolutionized structural biology by predicting the 3D shapes of proteins from their amino acid sequences. However, drug discovery is not merely a structural problem—it is a complex, multi-modal puzzle. A promising therapy must interact with biological targets, affect cellular pathways, avoid toxicity, and perform safely in human biology.[4]

To bridge these disparate domains, researchers at IBM have introduced MAMMAL (Molecular Aligned Multi-Modal Architecture and Language), a biomedical foundation model designed to treat diverse biological inputs as parts of a unified computational language. Published in the Nature portfolio journal npj Drug Discovery, MAMMAL represents a shift from specialized, single-task AI tools to a versatile, cross-modal architecture.[1]

The evidence supporting MAMMAL’s capabilities is anchored in its massive scale and diverse training data. The model was pre-trained on approximately 2 billion biological samples. This dataset spans four distinct modalities: protein sequences, antibody sequences, small-molecule representations, and gene expression profiles.[1][2][3]

By integrating these modalities, MAMMAL operates differently than large language models designed for human text. Instead of conversational prompts, researchers use a structured syntax to input molecular strings, amino acid sequences, or transcriptomic lab tests, allowing the model to learn the complex relationships across different biological domains.[1]

The MAMMAL architecture integrates four distinct biological modalities into a single foundation model.

The primary claim of the research is MAMMAL’s performance across a suite of standard drug discovery benchmarks. Evaluated on 11 diverse downstream tasks that span multiple stages of the pharmaceutical pipeline, the model achieved state-of-the-art (SOTA) results on nine of them, while remaining highly competitive on the remaining two.[1]

These benchmarks are not purely academic exercises; they represent critical hurdles in drug development. For instance, on molecular toxicity tests like ClinTox and blood-brain barrier penetration (BBBP), MAMMAL achieved Area Under the Receiver Operating Characteristic (AUROC) scores of 0.986 and 0.937, respectively. This represents a measurable improvement over previous leading models like MoLFormer.

MAMMAL achieved state-of-the-art results on 9 out of 11 evaluated drug discovery benchmarks.

The most heavily scrutinized claim in the evidence pack is MAMMAL’s performance relative to Google DeepMind’s AlphaFold 3. In a specific antibody-antigen binding benchmark, fine-tuned MAMMAL prediction scores were compared against AlphaFold 3’s confidence scores, which served as a proxy for binding likelihood.[3]

The data revealed that MAMMAL outperformed AlphaFold 3 in binding classification for five out of seven tested antigen targets. On larger, structurally complex targets like CD206 and VWF, MAMMAL demonstrated superior discriminative ability.[3]

The data revealed that MAMMAL outperformed AlphaFold 3 in binding classification for five out of seven tested antigen targets.

However, the researchers and independent analysts are careful to contextualize this finding. This result does not suggest that MAMMAL is universally superior to AlphaFold 3. AlphaFold 3 was explicitly designed for structural prediction and maintains an advantage on smaller targets where precise physical geometry is the primary driver of binding.

In specific antibody-antigen binding tests, MAMMAL outperformed AlphaFold 3 on five out of seven targets.

Instead, the evidence indicates that for specific classification tasks—where binding likelihood depends heavily on sequence context and cross-modal interactions—a modality-aligned foundation model can outperform a purely structural system. The consensus among analysts is that they are complementary tools rather than direct replacements.[4]

A significant strength of the MAMMAL project is its commitment to transparency and reproducibility. Unlike many commercial biomedical AI breakthroughs that rely on proprietary internal evaluations, IBM Research has made the 458-million-parameter model publicly available.[2]

The pretrained model weights are hosted on Hugging Face, and the fine-tuning codebase is accessible via the BiomedSciAI GitHub organization. This open-infrastructure approach allows independent researchers to reproduce the benchmark results, apply the model to proprietary datasets, and independently verify the AlphaFold 3 comparisons.[2]

For AI-native biotech startups and academic labs, this open access fundamentally alters the barrier to entry. Small teams working on antibody design or cancer drug response prediction can now leverage a massive foundation model without the prohibitive computational cost of pre-training from scratch.

Despite the strong benchmark performance, the evidence pack carries clear limitations and uncertainties. The most critical caveat is that computational benchmarks, no matter how rigorous, do not eliminate the need for physical experimentation.[4]

Open-source foundation models allow smaller biotech teams to accelerate early-stage drug discovery without massive computational budgets.

As one industry analysis noted, IBM Research's MAMMAL is not a miracle cure machine. The model can predict toxicity or binding affinity with high statistical accuracy, but drug candidates still fail in clinical trials at notoriously high rates because human biology is vastly more complex than any training dataset.

Furthermore, while MAMMAL integrates four major modalities, it does not yet capture the entirety of a living system's dynamic environment, such as real-time metabolic changes or complex immune system cascades. The predictions remain probabilistic hypotheses that require rigorous wet-lab validation.[4]

Looking forward, the introduction of MAMMAL signals a maturation in the field of AI-driven pharmacology. The bottleneck in drug discovery is increasingly not a lack of data, but the fragmentation of that data across isolated computational silos.

By proving that a single, unified architecture can process small molecules, proteins, and gene expression data simultaneously, MAMMAL provides a blueprint for the next generation of biomedical research. It moves the industry one step closer to an integrated computational ecosystem where the language of biology can be translated into viable therapeutics with unprecedented speed.[1][3]

Why this matters

Drug discovery is notoriously slow and expensive because biological data—from proteins to gene expression—is highly fragmented. By providing an open-source AI model that understands multiple biological 'languages' at once, researchers can accelerate the early stages of developing new medicines and treatments.

Sources

Source coverage

4 outlets

4 viewpoints surfaced

Computational Biologists 35%Open-Source Advocates 25%Structural Biologists 20%Clinical Translators 20%
  1. [1]arXivComputational Biologists

    MAMMAL -- Molecular Aligned Multi-Modal Architecture and Language

    Read on arXiv
  2. [2]Hugging FaceOpen-Source Advocates

    ibm/biomed.omics.bl.sm.ma-ted-458m

    Read on Hugging Face
  3. [3]ResearchGateComputational Biologists

    MAMMAL: A foundation model for cross-modal learning in drug discovery

    Read on ResearchGate
  4. [4]Factlen Editorial TeamComputational Biologists

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team

Comments

Stay informed

Every angle. Every day.

Get data analysis stories with full source coverage and perspective breakdowns delivered to your inbox.