Skip to main content

AI Medical Information Intelligence

What AI tells patients and HCPs about your product

We ask the questions your audience is already asking, to the models they are already using, and score every answer against your labeling — dose by dose, contraindication by contraindication, citation by citation. You get a report you can hand to medical, safety and legal in the same meeting.

From $250 · consulting request, not checkout

The gap

A medical information channel nobody is measuring

They ask the model first

A patient with a question about dosing at 11pm, or a prescriber checking an interaction between appointments, does not call your medical information line. They type it.

The answer leaves no trace

It appears in no inquiry record, no standard response document, and no safety feed. The answer was given in your product's name and you never saw it.

And it changes underneath you

Consumer interfaces route to models they do not name and change without notice. Whatever you checked by hand last quarter is not what is being returned today.

The audience line does not hold

Your labeling, your standard response documents and your congress materials were written for a healthcare professional, and the separation between that audience and a consumer is maintained by gates a model does not observe. A patient asks a plain question and gets back prescriber-grade content — dosing tables, titration, monitoring thresholds — synthesised from HCP-facing sources, delivered without the framing, the caveats or the clinician who was supposed to be standing between the two.

That is not hypothetical. It is visible in the published sample, where consumer interfaces returned regimen-level dosing to a plainly-worded patient question. Measuring it is the only way to know how far it goes for your product.

Sample report

Read the Report Sample

A sample report on tacrolimus immediate-release oral capsules: 20 questions put to 5 AI configurations, 100 responses captured verbatim, 16 findings adjudicated — 2 of them Critical. Nothing is gated, nothing is redacted, and nobody commissioned it.

What you get

The Medical Information AI Response Surveillance Report

One self-contained HTML document — it opens in any browser, including from a file you forward to someone with no account and no login — plus a PDF copy. Ten sections, in this order:

Executive summary

What was found, in four figures a non-specialist can act on: critical findings, high findings, dead citations, and citations that do not support their claim.

Dashboard

Findings by risk, recorded errors by configuration, critical-fact coverage, and citation behavior. Every figure has a stated formula in the methodology.

Model comparison

How the configurations differ, side by side — with the model label the interface actually displayed on the day, not a build string we cannot verify.

Domain analysis

Where the problems cluster by area of labeling: pregnancy and lactation, dosing, interactions, monitoring, off-label handling.

Key findings

The Critical and High findings in full, each linking to its detailed entry and the raw response behind it.

Detailed findings

Three layers kept strictly apart: what the model said, quoted verbatim; what the labeling says, quoted with its section; and our reading of the two.

Citations and sources

Every cited URL resolved and read against the statement it was offered for — not counted, read.

Model agreement

Where the configurations agree with each other and are wrong together, which is the failure mode a spot-check never catches.

Methodology and limitations

Every metric defined by formula, and a plain statement of what the report cannot tell you. A number a reader cannot reproduce is decoration.

Full appendix

Every question and every response it produced, captured verbatim and unedited, so any finding can be checked against its source.

How it works

Four steps, agreed before anything runs

01

Scope

We agree the product, the labeling revision it is measured against, the AI configurations, and who the questions speak as — a patient, a healthcare professional, or both. Nothing starts until the question corpus is agreed in writing.

02

Capture

Each question is asked once of each configuration, in a fresh conversation, in identical wording. No follow-ups. Every response is captured verbatim and unedited, with the date and the model label the interface displayed.

03

Score

Responses are scored against the labeling — critical facts present or absent, statements that contradict a named reference fact, and every citation resolved and read against the claim attached to it.

04

Adjudicate and release

Every finding is decided before release against the evidence behind it, and rejected findings are retained in the data rather than quietly dropped. You get the HTML report and a PDF copy.

Pricing

Published, so you can decide before you talk to anyone

Per product, per run. While the service is in build, every tier starts with a scoped consulting request rather than a checkout.

Snapshot
$250

one product, one point in time

The fastest way to find out whether there is a problem worth naming. Enough coverage to surface the obvious failures, and the full report format either way.

  • 20 questions, asked verbatim
  • Up to 3 major consumer-accessible AI providers
  • Full report: every finding, every response
Start a request →
Assessment
Recommended
$750

Enough questions to cover the labeling properly — special populations, interactions, monitoring, off-label handling — across the configurations your audience actually uses.

  • Approximately 50 questions
  • Up to 5 AI configurations, including reasoning modes
  • Domain analysis and model-agreement breakdown
Request the Assessment →
Program
Let's talk

recurring, across a portfolio

Consumer interfaces change without notice and without a version number. A single report is a photograph; a program is a measurement you can trend.

  • Multiple products, on a set cadence
  • Change tracked between runs, not re-stated
  • Scoped to what your review process needs
Start a request →

The published sample sits between the two priced tiers — 20 questions across 5 configurations — because it was built to show the report format rather than to fit a price. Read it and tell us which shape you actually need.

Scope

What this deliberately isn't

  • Not pharmacovigilance. Nothing here is an adverse event report, and it does not enter any safety system on your behalf.
  • Not a regulatory or labeling conclusion. The labeling is the reference standard, not the subject.
  • Not clinical advice for any individual patient.
  • Not a consistency study. One response per question per configuration — repeat runs are a separate exercise, and the report says so.
  • Not an audit of your medical information function. The subject is the model's answer, not your team's.
Questions

Before you ask

Is this an adverse event or safety reporting service?

No. This is an observational study of what consumer AI interfaces returned on recorded dates, measured against FDA labeling. Its recommendations are bounded to review, monitoring and governance considerations. Nothing in it is a pharmacovigilance, regulatory or labeling conclusion, and nothing in it is clinical advice for an individual patient. If a finding needs to enter your safety system, that is your process and your decision.

Which AI platforms do you cover?

The starting set is the interfaces your audience actually reaches: ChatGPT, Claude, Gemini, Google AI Overviews and OpenEvidence — the last of which matters because it is where a clinician goes deliberately for a clinical answer, rather than getting one incidentally from a general assistant. Providers and configurations are counted separately on purpose — two modes of the same product, such as standard versus extended reasoning, routinely return different answers, so each is scored as its own configuration rather than averaged together. A Snapshot covers up to three providers; an Assessment covers up to five configurations. The set is not fixed to the free consumer tiers: paid subscriptions, enterprise or in-house deployments, retrieval-backed assistants and any other surface your audience reaches can be added, and coverage expands as new interfaces appear. Which ones are worth paying to cover for your product is one of the things the consult decides.

How do I know the findings are real?

Every finding carries the full response it came from, quoted verbatim, alongside the labeling passage it is measured against and its section number — so you are never asked to take a judgement on trust. Every figure is defined by a formula in the methodology, and the appendix carries every question and every response it produced. The sample report is the whole thing, unredacted: read it and check one against the label yourself.

Why is it a consulting request instead of a checkout?

Because a report you did not help scope is a report you cannot use. The consult settles the things that decide whether the output is worth anything: the exact questions and their exact wording, who they are asked as, which models and configurations are searched, the labeling revision they are scored against, and what you need the finished report to survive — a safety review, a governance ask, a leadership conversation. All of that is agreed in writing before a single question runs. The prices are published so you can decide whether to start that conversation without booking a call first; the conversation exists so the deliverable matches what you actually needed.

Can I get it as a PDF?

Yes. The deliverable is a self-contained HTML report — it opens in any browser, including offline from a file you forward — and every report can be saved as a PDF from the report itself, with all the evidence expanded rather than collapsed behind a click. A PDF copy is supplied with every engagement.

Is the report confidential?

Your report is yours. The published sample is a specimen produced on our own initiative about a widely marketed generic, commissioned by nobody, and published to show the method. Nothing from a client engagement is published without written agreement.

How long does it take?

A Snapshot is normally a week from agreed scope; an Assessment is longer because the scoring and adjudication are the slow parts, and they are the parts worth having. A date is confirmed in the reply to your request, not before.

Request

Tell us the product and we'll scope it

A reply comes from a person, usually within a couple of days, with a proposed scope — the question corpus, the configurations, the labeling revision it will be measured against, and a date. No obligation and no booking link.

Add detail — optional

Your details are only used to reply to you.

By submitting, you agree to our Privacy Policy. Please don't include patient-identifiable information or adverse event details in this form.