Skip to content
universal/vqa

Medical Image Q&A

Ask any question about a medical image: draft answers with grounding boxes and prior comparison.

Coming soonANYUniversal & Multimodal

This model is not available to run yet. Vote to show interest and get an email when it opens.

Overview

What it does

Medical Image Q&A answers free-form questions about an image no dedicated MedRun model covers. Upload a few images, one CT or MR volume or a short clip and ask in plain language: "Is this a lateral decubitus film?", "What is the bright structure next to the gallbladder?". The bare model ID runs this question tool and returns an answer with a confidence and, where the image supports it, a box showing where the answer comes from.

The same run can ground a phrase ("the nodule", "right kidney") as bounding boxes, and with a prior study it describes what is new, resolved, larger or smaller. A separate tool writes a draft radiology-style report for any study, starting with a short structured description of what the images show.

Intended use

Orientation and drafting. A physician gets quick bearings on an unfamiliar study type; a developer explores what a study contains before choosing a dedicated model; an educator turns a teaching case into questions and answers. Every answer carries its provenance, and generated text never supplies numbers to a clinical report. When MedRun has a dedicated model for the study, the console suggests that model first.

Who it is for

Clinicians who want to ask about an image outside the dedicated catalogue, educators and residents exploring cases, and developers and researchers evaluating multimodal answers on their own data.

Inputs and protocol

Accepted input

  • Images: up to a few 2D images (radiographs, ultrasound frames, photographs, endoscopy stills) in DICOM, PNG or JPEG.
  • Volumes: one CT or MR series; slices are sampled to fit the model's context.
  • Clips: one short cine or video.
  • Question: free text; for grounding, the phrase to locate.
  • Prior (optional, in prior_files): an earlier image of the same region for the change description.

Requirements

Requirement Why
One study per question Answers refer to the images supplied, not to other series
Images in which the finding is visible after resizing Images are resized to the model's input size
Same region and comparable view for a prior The change description compares like with like

Optional context

A short history or indication can be passed in clinical_context; it is given to the model together with the question.

Outputs and standards

The result

Section Content
Answer Free-text answer with confidence and, where supported, a grounding box (the primary output of ask)
Study description Modality, region, view or sequence, technique and salient findings; returned by universal/vqa/report as the opening of its draft report
Phrase grounding Bounding boxes for a phrase such as "the nodule" or "right kidney"
Change versus prior Described change against the prior image: new, resolved, larger or smaller

Grounding boxes are returned as image coordinates, so every located phrase can be drawn on the source image.

Standards

  • JSON: report block flagged generative: true with provenance, citations[] for grounding boxes, changes[] for the prior comparison.
  • DICOM: SR TID 1500 planar regions for grounding boxes; GSPS overlays.
  • FHIR: DiagnosticReport with status preliminary for draft reports.

Related tools

  • universal/vqa/report: a draft radiology-style report (findings and impression) for any study, with the study description as its first section; chest radiographs and CT use specialised report weights. Returned as a FHIR DiagnosticReport with status preliminary.
  • universal/segment: outline the structure an answer refers to and measure it.

Result sections

One run returns every section its input supports.

  1. Study descriptiondescribe
  2. Phrase groundinglocate
  3. Change versus priorcompare
Explore all Universal & Multimodal models

Radiology Report Tools

Report text tools: narrative from findings, structuring, error check, coding, simplify, translate.

Whole body7 sections
ANY· text

Promptable Segmentation

Outline any structure in a medical image: click, box or scribble, name it, or label every organ.

Whole body1 sections
ANY· image-2d

Image Embeddings and Similar Cases

Medical image embeddings for search, clustering and zero-shot labels, plus similar-case retrieval.

Whole body0 sections
ANY· image-2d