Introducing MAVIS
Monitoring Assertions and Verifying Identifiers in Science
AI tools now draft a growing share of what research teams read: reports, literature reviews, summaries and data notes. MAVIS keeps a standing watch on the claims in those files. It has fred check each claim against public sources, checks again when something changes, and tells the person responsible when a claim cannot be traced, conflicts with its source, or stops holding.
Why MAVIS is needed
Pre-clinical decisions rest on claims: this compound binds that target, this paper showed that effect, this protein has that sequence. When people wrote every claim, a reviewer could ask where it came from. AI tools now write many of them, quickly, and they write wrong claims in the same confident style as right ones.
- AI writes faster than anyone can check.A literature review that took a week takes minutes. The checking has not sped up to match.
- AI invents references and identifiers that look real.Earlier studies found that 30 to 69% of LLM-generated references in biomedical writing were fabricated, and a 2026 test of 26 language models found reference generation unreliable across all of them.1,2 Fabricated references are now reaching the published record: an audit of 2.5 million biomedical papers found about one paper in 2,828 with a fabricated reference in 2023, and one in 277 in the first weeks of 2026.1
- Evidence moves.Papers are corrected and retracted, and database entries are revised. A claim that held when it was written can stop holding months later, while the document that carries it stays the same.
- Regulators expect traceable evidence, managed over time.The joint FDA and EMA principles of January 2026 cover AI used to generate or analyze evidence across the drug life cycle, including nonclinical work. They call for data governance and documentation, and for life cycle management of AI.3
A one-off check at the moment of writing is not enough. The claims a program relies on need a standing watch.
What MAVIS is
MAVIS is an application that sits beside a research program's files and watches the claims in them. It does not write content and it does not make decisions. It makes sure that every claim it can check has been checked, keeps checking it, and shows people the result with its evidence.
- It watches your files
- Your program's notes, reports and data, as Markdown and JSON files with a README that says what each file is and where it came from. Files drafted by AI are marked as such.
- It has fred check every claim
- Any statement that names an identifier (a PubMed ID, DOI, UniProt or PubChem entry) or a value with a unit is sent to fred, which looks it up in public databases.
- It alerts the owner
- When a claim cannot be traced, disagrees with its source, or stops holding, the person responsible for it is told, once, with the reason.
- It shows the provenance
- For every claim: the file and line it came from, the check that tested it, each database lookup behind the result, and the record it found.
People stay in charge. The claim's owner reads the alert and its evidence, then corrects the file, accepts the claim with a note, or withdraws it. MAVIS records the decision.
How it works
MAVIS runs continuously, in a loop of four steps.
- Probe the files. MAVIS looks for new or changed files and picks out every claim in them that names an identifier or a value with a unit.
- Ask fred. It sends the claims to fred, which checks each one against public databases. A claim never counts as its own evidence: MAVIS does not send a claim's own file as the material to check it against.
- Read fred's result. For each claim, the verdict and every lookup behind it.
- Compare and alert. MAVIS compares the result with the last check. A claim counts as holding only when a public source confirmed it. A change raises an alert only when a second check agrees, because AI-based checks can vary from one run to the next.
When it checks
When a file is added or edited, when a scheduled re-check is due, and when a public source may have changed, such as a paper being retracted.
What the owner sees
- Cannot be traced. No database returns the identifier or citation.
- Conflicts with its source. The claim disagrees with the database record.
- Evidence has changed. A cited paper was retracted or corrected, or a claim that held no longer does.
- Unsupported. An AI-drafted file makes a claim with nothing in it that can be checked.
- Still holds. No alert. The check is logged quietly on the claim.
Everything appears on the MAVIS dashboard: what was checked and when, open alerts, claims labeled by where their support comes from, and the provenance of each claim.
How MAVIS uses fred
fred is a reasoning engine for pre-clinical drug discovery that runs on your own machine with a free open-weight model. It looks things up in public databases and traces every identifier in its answers to where it came from. MAVIS does not change fred. It calls fred the way any other application does, and adds the watching, the memory of past checks, the alerts and the dashboard.
fred does
- Reasons with a local open-weight model.
- Looks values up in UniProt, PubChem, PubMed, Open Targets, Reactome and the web.
- Traces each identifier to the lookup or file it came from, and flags what it cannot trace.
- Keeps a record of every run.
MAVIS does
- Watches your files and finds the claims in them.
- Decides what to check and when, and asks fred.
- Remembers every past check, so it can see change.
- Raises alerts, shows the dashboard and records decisions.
MAVIS gives fred your program's own Markdown and JSON files, with a README that labels each one, and asks fred for a verdict on every claim: confirmed by a public source, in conflict with it, not found, or supported only by your own files. fred also reports when a paper a claim cites has been retracted or corrected.
Because both run on your own machine, your files and fred's reasoning never go to a model provider. This is private, not air-gapped: fred's searches of public databases do leave the machine.
What MAVIS checks, and what it does not
MAVIS checks whether each claim's identifiers and values trace to a public source and still agree with it. It does not check whether the science is right; that judgment stays with people.
- It only checks claims that name an identifier or a value with a unit. Statements with neither are counted and shown as outside its coverage, never as checked.
- A claim that rests on your own experimental data cannot be confirmed by a public database. MAVIS labels it as internal rather than leaving it permanently unverified.
- A real identifier attached to the wrong thing is the hardest error to catch, and MAVIS will not catch every one.
- AI-based checks can vary between runs, so MAVIS can raise false alarms. Owners can mark them, and MAVIS counts them.
- A check that fails is shown as a failure. It is never shown as a pass.
Sources
- Topaz M, et al. Fabricated citations: an audit across 2.5 million biomedical papers. The Lancet 407 (2026): 1779-81. Summary: casrai.org.
- Biomedical reference generation remains unreliable across 26 large language models. arXiv 2609.14988 (2026): arxiv.org/pdf/2609.14988.
- FDA and EMA. Guiding principles of good AI practice in drug development, January 2026: fda.gov/media/189581/download.