AI clinical-data platform

Clinical data, module by module.

MedCupola is a light core with plug-in modules: import any format, ask questions in plain English, and keep every installation validated — on your own servers.

  • Migrate from anything
  • Private AI
  • Validated by design
MedCupola architecture: users, kernel, modules, platform services and the three-step ingestion pipeline

The problem

Clinical data is scattered, and the systems are rigid.

Dozens of formats

Data arrives as CSV, Excel, XML, EDC exports, CDISC, FHIR, or a database dump — and every new source means a new, brittle import.

Monolithic tools

Most platforms are all-or-nothing. You buy the whole thing and live with the parts you don't need.

AI bolted on

“AI” is often a chat box on the side, disconnected from the data, with no guardrails and no evidence.

Validation last

Proving a system is correct becomes a painful, one-time project instead of something that simply stays true.

The idea

Built like a space station.

There is a central core that provides the essentials — power, life support, a docking ring. Over the years, modules are added or removed: a laboratory here, a solar array there. Each attaches to a standard interface, and the station keeps working throughout.

The core

A small, domain-agnostic kernel: accounts, storage, a central record of every action, a model gateway, and the ability to host add-ons.

  • registry
  • lifecycle
  • event bus
  • config
  • audit
  • storage
  • model gateway
  • MCP host
  • UI shell

The modules

Each capability is a pip-installable package that plugs in, activates, and can be swapped — independently of everything else.

  • clinical-core
  • validation
  • sql-agent
  • rag
  • assistant
  • data-quality
  • mapping
  • ingestion adapters

Migrate from anything

Any format in, in three steps.

Point MedCupola at almost any source. The format is detected, the mapping is proposed, and the data is loaded, checked, and kept validated from then on.

  1. 1

    Bring anything

    CSV, Excel, XML, JSON, EDC exports, CDISC, FHIR, or a database dump. The format is detected automatically.

  2. 2

    Confirm the mapping

    The AI proposes how each source field maps to the target schema. A person reviews and approves.

  3. 3

    Load and validate

    The data is loaded, checked, and reported — and every subsequent change is revalidated automatically.

Any format in  ·  one shape inside  ·  every migration validated

AI-first

AI is the product, not a feature.

The agent suite ships in the core, not as a later add-on. Agents plan, use tools over a standard interface, and answer with evidence — under guardrails that keep them read-only unless a person approves.

Ask in plain English

The SQL agent turns a question into a safe, read-only query, runs it, and shows its work.

Answers with citations

Document search grounds every response in the exact source, so nothing is invented.

Automatic data checks

Data-quality and mapping agents flag missing, invalid, duplicate, or inconsistent records with evidence.

  • agent-runtime
  • sql-agent
  • rag
  • assistant
  • data-quality-agent
  • mapping-agent
  • guardrails
  • prompt-eval

Modules

Start small, add modules as you grow.

Run and pay for only what you need, then plug in more. New formats are new adapters; new clinical needs are new modules.

Tier 1 — the foundation

clinical-core

Synthetic study schema, seed data, and vector search (pgvector).

validation

IQ/OQ/PQ, perpetual revalidation, traceability, and gating.

agent suite

Runtime, SQL, RAG, assistant, data-quality, mapping, guardrails, eval.

Tier 2 — first add-ons

adjudication (CEC/EAC)

Blinded endpoint and event review: intake, dossier assembly, committee voting, consensus, and export.

xml-ingestion

EDC XML → validate → map → load: the first proof of pluggability.

core-lab · data-quality · mapping

Lab import with reference ranges; deterministic checks; human-approved schema mapping.

Tier 3 — beyond the pilot

The catalog runs much deeper. As the platform grows, these plug in exactly the same way — a new format or a new clinical need is just a new module, with no change to the core.

  • EDC
  • ePRO / eCOA
  • eConsent
  • CTMS
  • eTMF
  • RTSM / IWRS
  • safety & pharmacovigilance
  • CDISC (SDTM / ADaM)
  • MedDRA · WHO Drug · LOINC coding
  • real-world evidence
  • biostatistics & TLF
  • imaging review
  • risk-based monitoring
  • eSubmission (eCTD)

Trust

Validated by design, private by default.

Validation is a module

A dedicated module confirms each installation is correct and re-checks the entire system whenever anything changes.

  • IQ / OQ / PQ qualification
  • Perpetual revalidation on upgrade
  • Traceability and evidence
  • Bad upgrades blocked and quarantined

Your data never leaves your servers

Inference runs locally. A small model is always on; extra GPU power is switched on only when needed.

  • Local-only models — no external AI calls
  • On-demand GPU, off when idle
  • Read-only guardrails and approvals
  • Synthetic-data demos, safe to share

The pilot

Run a pilot with us.

We are running a focused pilot on synthetic data — deploying the platform, migrating real-shaped datasets through the 1-2-3 pipeline, and proving the AI agents end to end. We are looking for a partner to shape it.

Request a pilot