WHAT AI CAN'T LEARN FROM THE INTERNET
- BioPrime Stories & latest news
- WHAT AI CAN'T LEARN FROM THE INTERNET
WHAT AI CAN'T LEARN FROM THE INTERNET

As AI becomes pervasive across agriculture, the advantage will increasingly come from access to information others simply do not have
There's something counterintuitive buried in that sentence. AI is supposed to be the great leveler — the technology that puts sophisticated capability into anyone's hands, regardless of budget or headcount. And in one sense, it genuinely does: a small team today can access language models, prediction engines, and analytical tools that would have required a research division a decade ago.
But pull the lens back and a different pattern appears. The more powerful and widely available AI becomes, the more the remaining advantage concentrates in a narrower place: not in who has the best model, but in who has fed it something nobody else has. Information, ownership of that information, and the legal freedom to actually use it are becoming the real moat — and almost nowhere is that clearer than in biological discovery.
The Debate AI Investors Are Still Having
This isn't a settled question, and it's worth being honest about that. In 2019, Andreessen Horowitz published an influential essay, “The Empty Promise of Data Moats,” arguing that proprietary data is a far weaker competitive advantage than founders claim. Their reasoning: data network effects weaken quickly past a certain scale, competitors can often collect similar data themselves, and — the sharpest point — the value of data depends entirely on the specific task at hand. General-purpose data doesn't help with a specialized task, no matter how much of it you have.
Five years later, the same firm revisited the question and reached a more nuanced conclusion. In “The Race to Capture Value: Cloud Lessons for the AI Era” (2024), they wrote that data does provide a durable advantage “when AI products depend on proprietary data or data scale as a crucial ingredient and key differentiator.” Their example: KoBold Metals, a mineral exploration company that built exclusive commercial agreements for historical geological survey data — records that took decades of physical drilling and surveying to generate, and that no AI model could conjure from public sources.
The resolution to the apparent contradiction is simple once you see it: generic data is a weak moat; proprietary, hard-to-replicate, physically-generated data tied to a specific decision is a strong one. Almost every industry sits somewhere on that spectrum. Biology sits close to the strong end — and it's worth understanding exactly why.
What AlphaFold Proves — and Where It Stops
In 2024, Demis Hassabis and John Jumper won the Nobel Prize in Chemistry, alongside David Baker, for AlphaFold — an AI system that predicts a protein's three-dimensional structure directly from its amino acid sequence, a problem that had resisted direct solution for fifty years. It is one of the most genuine AI-driven scientific breakthroughs on record, and by October 2024 it had been used by more than two million researchers in 190 countries.
Here is the detail worth sitting with: AlphaFold was only possible because of decades of publicly deposited, freely shared experimental data in the Protein Data Bank — structures painstakingly solved by thousands of crystallography labs over half a century and made openly available. AlphaFold didn't replace that data. It was built entirely on top of it, and its output — predicted structures for virtually every known protein — is now itself freely public.
Structure prediction became solvable, and then free, because the underlying data was already public. Function, in a specific place, under specific conditions, is a different problem entirely.
Knowing what shape a protein folds into is enormously useful, but it does not tell you whether a specific microbial strain will solubilize iron-bound phosphorus in an acidic Deccan soil during a monsoon-disrupted growing season. Nobody has published that dataset, because nobody has generated it at scale — it requires exactly the kind of slow, expensive, physical experimentation that AI cannot substitute for, only interpret once it exists.
Decoding “Proprietary Information”: Access, Ownership, and the Right to Use It
The phrase gets used loosely, but it actually breaks into three distinct, related ideas.
Access
Some data can be scraped, licensed, or purchased — which means, by definition, a competitor can eventually get it too. Other data can only be generated through direct physical or experimental work: running an assay, screening a strain against a specific stress, completing a multi-season field trial. That second category is what economists would call genuinely scarce — its supply doesn't expand just because demand for it does.
Ownership: Patents vs. Trade Secrets
These are the two main legal containers for proprietary knowledge, and they behave very differently:

A large, continuously growing proprietary dataset is often better protected as a trade secret than as a patent, precisely because a patent eventually expires and becomes public — while a well-guarded dataset that keeps generating new discoveries never does.
Freedom to Operate
This is the quieter, less-discussed piece, and AI has made it sharply relevant. Freedom to operate (FTO) is the ability to actually commercialize a discovery without infringing someone else's existing patents. An AI model trained heavily on public chemical or biological databases runs a real risk of “rediscovering” compounds or structures that are already covered by someone else's prior art — which is a problem you only find out about after significant investment, not before.
A Live Example: Why a Public Drug-Discovery Company Bet Everything on Proprietary Data
Recursion Pharmaceuticals (Nasdaq: RXRX) is a useful case study precisely because it operates in public markets and has to explain its strategy plainly. Over the past decade it has built what it describes as one of the largest fit-for-purpose proprietary biological and chemical datasets in the world — more than 50 petabytes spanning phenomics, transcriptomics, proteomics, and more, generated by an automated wet lab running millions of cell experiments a week.
Industry analysts have been explicit about why this matters beyond scale: “proprietary data generation is the deepest form of IP moat in AI drug discovery: it is harder to circumvent than a patent and does not expire.” And the freedom-to-operate connection is direct — by training its models primarily on data it generated itself rather than public databases, Recursion reduces the risk of its AI proposing compounds that fall inside someone else's existing patent space. Proprietary data isn't just a research asset here. It's a legal risk-management strategy.
The Line That Keeps Moving

As AI matures, generic capability sinks below the line and becomes freely available; advantage concentrates in what remains above it.
This is the pattern underneath everything above. As AI capability spreads, whatever can be learned from public information gets commoditized — cheaper, faster, and available to everyone at once. What's left as genuine advantage is exactly what public information never contained in the first place: proprietary, physically-generated, context-specific knowledge, held as a trade secret, and used without creating freedom-to-operate risk.
Why This Holds True for BioPrime
This is precisely the position BioNexus and SNIPR occupy. BioNexus is not a public database — it's one of India's largest plant-associated microbiome libraries, over 18,000 strains, built through years of field collection across Indian soils and cropping systems that no public repository documents. SNIPR is the discovery engine that turns that library into structured, functional trait data: which strain solubilizes which form of phosphorus, under which pH, against which stress, on which crop.
None of that dataset exists anywhere on the public internet. It cannot be scraped, purchased, or reconstructed by a competitor's AI model, because it was never published — it was generated, one screen and one field trial at a time, and it continues to compound with every new experiment run. That makes it, in the language of this piece, exactly the kind of proprietary information that sits above the commoditization line: physically generated, context-specific, held largely as trade secret rather than public disclosure, and — because it wasn't trained on anyone else's patented compounds — clean from a freedom-to-operate standpoint.
As AI becomes a standard layer in agricultural discovery — as it inevitably will — a generic model trained on public genomic databases will be able to tell you a great deal about what a microbe could plausibly do. It will not be able to tell you, with any real confidence, whether a specific strain from BioNexus reliably mobilizes iron-bound phosphorus in a Vertisol under Maharashtra's monsoon conditions. That answer only exists because BioPrime went and generated it — and that gap, multiplied across thousands of strains and hundreds of crop-stress combinations, is the moat.
The AI layer will be built by many. The information it needs to be trustworthy, in a specific field, for a specific crop, will not be available to all of them.
The Advantage That Compounds
None of this is an argument against AI, or against the wave of prediction and discovery tools now reaching agriculture. It's an argument about where the durable advantage will actually sit once those tools are common infrastructure rather than novelty. Algorithms will keep getting better, cheaper, and more widely available — that trend is not reversible, and shouldn't be resisted.
What won't become common, no matter how good AI gets, is the proprietary, physically-generated, field-validated information that AI needs in order to be right rather than merely plausible in a specific place, on a specific crop, under a specific stress. Owning that information — legally, structurally, and through the years of costly work required to generate it — is what turns a good AI model into a trustworthy one, and it is the single clearest source of durable advantage left standing as everything else gets commoditized around it.
References & Further Reading
- Andreessen Horowitz — The Empty Promise of Data Moats (2019) — https://a16z.com/the-empty-promise-of-data-moats/
- Andreessen Horowitz — The Race to Capture Value: Cloud Lessons for the AI Era (2024) — https://a16z.com/cloud-lessons-for-the-ai-era/
- Nature — Chemistry Nobel Goes to Developers of AlphaFold AI That Predicts Protein Structures — https://www.nature.com/articles/d41586-024-03214-7
- The Nobel Prize — Press Release: The Nobel Prize in Chemistry 2024 — https://www.nobelprize.org/prizes/chemistry/2024/press-release/
- The Nobel Prize — Popular Information: The Nobel Prize in Chemistry 2024 — https://www.nobelprize.org/prizes/chemistry/2024/popular-information/
- Recursion Pharmaceuticals — Pioneering AI Drug Discovery (company platform overview) — https://www.recursion.com/
- DrugPatentWatch — Patent Data Is the Missing Ingredient Powering AI Drug Discovery — https://www.drugpatentwatch.com/blog/the-missing-ingredient-why-patent-data-is-the-key-to-unlocking-ai-powered-drug-discovery/
- Mawer Investment Management — Data Moats in the Age of AI: What Still Matters? — https://www.mawer.com/the-art-of-boring/blog/data-moats-in-the-age-of-ai-what-still-matters
Note: the data-moat debate referenced here is genuinely live among technology investors and strategists, and reasonable analysts land on different sides of it depending on the industry and data type in question; this piece takes the position that the debate itself resolves cleanly once “data” is separated into generic/public versus proprietary/physically-generated categories, which is the framing argued for above rather than a settled consensus view.
