LLM applications and healthcare data
Using large language models with healthcare data: the context API request and response, cloud and local models, agentic systems, retrieval versus context, and prompting with provenance.
An application that uses a large language model with healthcare data has to solve one problem the model cannot: which part of a patient's record to put in the prompt, and how to label it so the model can tell facts from notes. Anpheros solves it on the data side with a context API: given a patient, a task and a token budget, it returns the relevant sections of the patient's HL7 FHIR R4 record, each item labelled with its source. Your application sends that context to the model you choose — hosted or local.
How do I connect an LLM to FHIR medical data? Keep the records in a FHIR store the model cannot reach directly — with Anpheros, the patient's HL7 FHIR R4 record — ask it for a budgeted, source-labelled context (POST /v1/context), and put that context in the prompt of the model you choose. The model never receives database or API credentials; your application does the calls, within the patient's consent.
Anpheros does not run models and has no native integration with any model provider or local runtime.
The pattern
1. your app ── POST /v1/context {patient, task, question, budget_tokens, format} ──► Anpheros
2. Anpheros ── sections + text + omitted + warnings + manifest_id ──────────────────► your app
3. your app ── system prompt + context + user question ─────────────────────────────► the model
4. the model ── answer ───────────────────────────────────────────────────────────────► your app
5. your app ── optional write-back with author_type "ai" ──────────────────────────► Anpheros
The context request
curl -X POST https://platform.anpheros.com/v1/context -H "Authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{
"patient": "'$PID'",
"task": "medication review before a cardiology visit",
"question": "How did blood pressure evolve over the last 3 months?",
"budget_tokens": 1500,
"format": "text"
}'
| Field | Meaning |
|---|---|
patient |
the patient id your credential sees |
task, question |
what the model will do; the platform plans which parts of the record are relevant |
needs |
optional explicit needs, for example ["labs:4548-4", "vitals:trend:85354-9", "timeline:180d"] |
budget_tokens |
size of the context, 300–8 000 (default 2 000) |
format |
structured (JSON sections) or text (also returns a ready-to-use text block) |
window_days |
optional look-back window |
The context response
sections— for example a summary card, conditions, medications, allergies, immunizations, lab results, vital signs with weekly trends and before/after-treatment markers, symptoms, timeline and documents. Every item carriesauthor_typeandsource.omitted— what did not fit the budget; tell the model its context is partial.warnings,provenance_note.text— whenformatistext.manifest_id— the record of which sections and sources were used, tied to the access log.- an
ai_notessection — values that an AI wrote earlier appear only there, marked as not verified, never mixed with the clinical sections.
Cloud models
With a hosted model — from OpenAI, Anthropic, Google, xAI or another provider — your backend puts the context in the prompt and calls the provider's API. The context leaves your infrastructure for that provider, so:
- make sure the patient's consent and your privacy notice cover sending data to it;
- have a data processing agreement with the provider that fits health data;
- send only what the task needs — the token budget and explicit
needshelp.
Local models
With a model you run yourself — for example through Ollama or another local runtime — the only network call carrying patient data is the one between your backend and Anpheros; the prompt and the answer stay on your infrastructure. Smaller local models have smaller context windows: lower budget_tokens accordingly and prefer format: "text".
Agentic systems
When the model decides which calls to make, give it tools (context, search, timeline) instead of raw credentials, keep scopes narrow and log each manifest_id next to the answer. AI agents and medical data
Context API or your own retrieval?
A retrieval pipeline over free text (embedding chunks of documents and searching them) is useful for unstructured notes. For structured medical data the context API is usually simpler and safer: it works on coded FHIR resources, keeps provenance, knows about trends and treatment periods, respects consent and is audited. The two can be combined — for example context from Anpheros plus retrieved passages from documents you manage.
Prompting with provenance
Keep the labels in the prompt and tell the model what they mean:
The context below comes from the patient's record. Each item is labelled with its author type
(patient, practitioner, device, import, derived) and source. Items under ai_notes were written by an
AI earlier and are not verified. Say which items your answer relies on. If the context is marked as
partial, say so. Do not give a diagnosis or change a treatment; suggest discussing it with a clinician.