Data

Image abstractions, patches, annotations, and regions of interest.

Image (Abstract Base)

class medical_image.data.image.Image[source]

Bases: ABC

Abstract base class for medical images.

Supports lazy loading and four mutually-exclusive construction paths (file, array, source image, or empty shell). Width and height are computed properties derived from pixel_data.shape when loaded, falling back to cached values before loading.

The Image optionally holds a list of Annotation objects via aggregation – an image can exist without annotations.

file_path

Path to the image file on disk.

Type:

Optional[str]

pixel_data

Pixel values (None until loaded).

Type:

Optional[torch.Tensor]

annotations

Attached annotations (None by default).

Type:

Optional[List[Annotation]]

__init__(file_path=None, array=None, width=None, height=None, source_image=None)[source]

Initialise an Image via one of four construction paths.

Parameters:
  • file_path (str | None) – Path to an image file. Raises FileNotFoundError if it does not exist.

  • array (ndarray | Tensor | None) – Pre-existing numpy array or torch tensor to wrap as pixel data.

  • width (int | None) – Explicit width hint (used before pixel data is loaded).

  • height (int | None) – Explicit height hint (used before pixel data is loaded).

  • source_image (Image | None) – Another Image to clone metadata and pixel data from.

property width: int | None
property height: int | None
property device: device
to(device)[source]

Move pixel data to device (in-place).

Parameters:

device (str | device) – Target device (e.g. "cuda", "cpu").

Returns:

self, for method chaining.

Return type:

Image

ensure_loaded()[source]

Raise DicomDataNotLoadedError if pixel data has not been loaded.

Returns:

self, for method chaining.

Return type:

Image

pin_memory()[source]

Pin pixel data to page-locked memory for faster GPU transfers.

No-op if pixel data is None or already pinned.

Returns:

self, for method chaining.

Return type:

Image

clone()[source]

Create a lightweight copy of this image.

Clones the pixel data tensor and shallow-copies the annotation list, but does not copy heavy objects (DICOM dataset, PIL image).

Returns:

A new Image of the same concrete type.

Return type:

Image

classmethod from_file(file_path)[source]

Construct an Image from a file path (lazy – does not load pixels).

Parameters:

file_path (str)

Return type:

Image

classmethod from_image(other_image)[source]

Construct an Image by copying metadata and pixel data from other_image.

Parameters:

other_image (Image)

Return type:

Image

classmethod from_array(array)[source]

Construct an Image from a NumPy array or PyTorch tensor.

Parameters:

array (ndarray | Tensor)

Return type:

Image

classmethod empty(width=None, height=None)[source]

Construct an empty Image shell with optional width/height hints.

Parameters:
  • width (int | None)

  • height (int | None)

Return type:

Image

abstractmethod load()[source]

Load pixel data (lazy load). Must be implemented by subclasses.

abstractmethod save()[source]

Save pixel data. Must be implemented by subclasses.

add_annotation(annotation)[source]

Append an annotation to this image.

Initialises the annotation list to [] on first call if it is currently None.

Parameters:

annotation (Annotation) – The Annotation to attach.

Return type:

None

remove_annotation(index)[source]

Remove and return the annotation at index.

Parameters:

index (int) – Zero-based position in the annotation list.

Returns:

The removed Annotation.

Raises:

IndexError – If the annotation list is None or index is out of range.

Return type:

Annotation

to_json(file_path=None)[source]

Serialize this image’s metadata and annotations to JSON.

Pixel data is not included – only file path, dimensions, image type, and the full annotation list.

Parameters:

file_path (str | None) – If provided, the JSON string is also written to this file path.

Returns:

A JSON string with keys file_path, width, height, image_type, and annotations.

Return type:

str

classmethod from_json(json_input)[source]

Deserialize an Image from a JSON string or file path.

Pixel data is not loaded – only metadata and annotations are restored. Call on a concrete subclass (InMemoryImage, DicomImage, PNGImage). For automatic subclass dispatch see image_from_json().

Parameters:

json_input (str) – A JSON string or a path to a .json file.

Returns:

A new Image instance with annotations attached.

Return type:

Image

display_info()[source]

Log summary information about this image (path, dimensions, device, annotations).

Return type:

None

medical_image.data.image.requires_loaded(func)[source]

Decorator that verifies pixel_data is loaded on every Image argument.

Inspects all positional and keyword arguments; raises DicomDataNotLoadedError if any Image has pixel_data is None.

DicomImage

class medical_image.data.dicom_image.DicomImage[source]

Bases: Image

DICOM image backed by pydicom.

Supports lazy loading: the constructor stores the file path and validates the extension; pixel data is read only when load() is called.

dicom_data

The parsed DICOM dataset (None until load() is called).

Type:

Optional[pydicom.Dataset]

__init__(file_path=None, array=None, width=None, height=None, source_image=None)[source]

Initialise a DICOM image.

Parameters:
  • file_path (str | None) – Path to a .dcm file.

  • array (ndarray | Tensor | None) – Pre-existing pixel data (numpy or tensor).

  • width (int | None) – Explicit width hint.

  • height (int | None) – Explicit height hint.

  • source_image (Image | None) – Another Image to clone from.

Raises:

ValueError – If file_path does not have a .dcm extension.

load()[source]

Read the DICOM file and populate pixel_data, width, and height.

Return type:

None

save()[source]

Write modified pixel data back to {name}_modified.dcm.

Raises:

ValueError – If dicom_data has not been loaded.

Return type:

None

PNGImage

class medical_image.data.png_image.PNGImage[source]

Bases: Image

PNG image backed by Pillow.

Supports lazy loading: the constructor validates the file extension; pixel data is read only when load() is called.

_pil_image

The Pillow image object (None until load() is called).

Type:

Optional[PIL.Image.Image]

__init__(file_path)[source]

Initialise a PNG image.

Parameters:

file_path (str) – Path to a .png file.

Raises:

ValueError – If file_path does not have a .png extension.

load()[source]

Open the PNG file via Pillow and populate pixel_data as a float tensor.

Return type:

None

save()[source]

Write pixel data to {name}_modified.png as uint8.

Raises:

DicomDataNotLoadedError – If pixel_data is None.

Return type:

None

InMemoryImage

class medical_image.data.in_memory_image.InMemoryImage[source]

Bases: Image

Concrete image that lives only in memory (no file I/O).

Both load() and save() are no-ops. Useful for intermediate processing results, temporary images, and test fixtures.

__init__(file_path=None, array=None, width=None, height=None, source_image=None)[source]

Initialise an in-memory image.

Accepts the same parameters as Image.

Parameters:
load()[source]

No-op (in-memory images have no file to load).

Return type:

None

save()[source]

No-op (in-memory images have no file to save to).

Return type:

None

PatchGrid

class medical_image.data.patch.PatchGrid[source]

Bases: object

Divides an Image into a regular grid of rectangular patches.

Automatically pads the image with zeros when its dimensions are not evenly divisible by the requested patch size.

parent

The source image.

Type:

Image

patch_h

Height of each patch in pixels.

Type:

int

patch_w

Width of each patch in pixels.

Type:

int

patches

Flat list of all patches (row-major order).

Type:

list[Patch]

grid

2-D grid indexed as grid[row][col].

Type:

list[list[Patch]]

pad_bottom

Number of rows of zero-padding added at the bottom.

Type:

int

pad_right

Number of columns of zero-padding added on the right.

Type:

int

__init__(parent_image, patch_size)[source]

Create a PatchGrid by splitting parent_image.

Parameters:
  • parent_image (Image) – The image to divide. Must have pixel_data loaded (i.e. not None).

  • patch_size (Tuple[int, int] | int) – (patch_height, patch_width) or a single int for square patches.

classmethod from_image(image, patch_size)[source]

Create a PatchGrid from an Image.

Loads the image lazily if its pixel_data is None.

Parameters:
  • image (Image) – Source image.

  • patch_size (Tuple[int, int] | int) – (patch_height, patch_width) or a single int for square patches.

Returns:

A new PatchGrid covering the entire image.

Return type:

PatchGrid

reconstruct()[source]

Reassemble the full image tensor from patches (removing padding).

Uses pre-allocated output tensor for O(1) allocations instead of O(rows * cols) concatenations.

Returns:

The reconstructed torch.Tensor with the same layout as the original image.

Return type:

Tensor

to_image()[source]

Reconstruct the full image from patches and return a new Image.

Removes any padding that was added during splitting, clones the parent image, and replaces its pixel data with the reconstructed tensor. Follows the same pattern as load().

Returns:

A new Image containing the reassembled pixel data.

Return type:

Image

Patch

class medical_image.data.patch.Patch[source]

Bases: object

Represents a patch extracted from an image.

parent

The parent image from which this patch is extracted.

Type:

Image

row_idx

Vertical index in the patch grid.

Type:

int

col_idx

Horizontal index in the patch grid.

Type:

int

row_offset

Top-left row (height) pixel coordinate.

Type:

int

col_offset

Top-left column (width) pixel coordinate.

Type:

int

pixel_data

Tensor representing patch pixels.

Type:

torch.Tensor

is_padded

Whether the patch contains padding.

Type:

bool

__init__(parent, row_idx, col_idx, x, y, pixel_data, is_padded=False)[source]
Parameters:
property x: int

Row offset (kept for backward compatibility).

property y: int

Column offset (kept for backward compatibility).

property height: int
property width: int
grid_id()[source]

Return patch position as (row, col) in the grid.

Return type:

Tuple[int, int]

pixel_position()[source]

Return top-left (x, y) pixel coordinates in the original image.

Return type:

Tuple[int, int]

to_numpy()[source]

Convert patch pixel data to a NumPy array.

Return type:

ndarray

to_image()[source]

Convert this patch into a new Image instance.

Clones the parent image and replaces its pixel data with the patch tensor. Follows the same pattern as load().

Returns:

A new Image object containing only the patch pixels.

Return type:

Image

load()[source]

Convert this patch into a new Image instance.

Deprecated since version Use: to_image() instead for consistency with the ROI API.

Returns:

A new Image object containing only the patch pixels.

Return type:

Image

RegionOfInterest

class medical_image.data.region_of_interest.RegionOfInterest[source]

Bases: object

PyTorch-compatible Region of Interest (ROI) extractor.

Crops a sub-region from an Image using one of three coordinate formats:

  • Bounding Box: [x_min, y_min, x_max, y_max]

  • Polygon: [(x1, y1), ..., (xn, yn)]

  • Mask: 2D boolean NumPy array

image

The source image.

Type:

Image

coordinates

ROI definition (format depends on annotation type).

annotation_type

Detected ROI type.

Type:

GeometryType

__init__(image, coordinates)[source]

Initialise an ROI.

Parameters:
  • image (Image) – Source image (will be loaded lazily if needed).

  • coordinates (List[int] | List[Tuple[int, int]] | ndarray) – ROI definition – bounding box, polygon, or mask.

classmethod from_center(image, cx, cy, half_size)[source]

Create a bounding-box ROI from center coordinates and half-size.

Parameters:
  • image (Image) – Source Image.

  • cx (int) – Center row (y-axis in image space).

  • cy (int) – Center column (x-axis in image space).

  • half_size (int) – Half-size of the square ROI.

Returns:

RegionOfInterest with bounding box coordinates.

Return type:

RegionOfInterest

load()[source]

Crop the image using the ROI definition and return a new Image.

Loads the source image lazily if it has not been loaded yet.

Returns:

A cloned Image whose pixel_data contains only the cropped region.

Return type:

Image

static normalize(image, divisor=4095.0)[source]

Normalize pixel values by dividing by a constant (e.g. 4095 for 12-bit).

Modifies the image in-place and returns it.

Parameters:
  • image (Image) – Image to normalize.

  • divisor (float) – Value to divide by.

Returns:

The same Image with normalized pixel_data.

Return type:

Image

Annotation

class medical_image.data.annotation.Annotation[source]

Bases: object

A single annotation on a medical image.

Represents a geometric region (rectangle, ellipse, or polygon) with a label and optional metadata. The centroid is computed automatically in the constructor and exposed as center.

shape

Geometry type of the annotation.

Type:

GeometryType

coordinates

Shape-specific coordinate data.

label

Human-readable annotation label.

Type:

str

metadata

Arbitrary extra information (BI-RADS, pathology, …).

Type:

dict

center

Computed centroid (cx, cy).

Type:

Tuple[float, float]

__init__(shape, coordinates, label, metadata=None)[source]

Initialise an annotation and compute its center.

Parameters:
  • shape (GeometryType) – Geometry type (RECTANGLE, ELLIPSE, or POLYGON).

  • coordinates (List[int] | List[Tuple[int, int]]) – Coordinate data whose format depends on shape: - RECTANGLE: [x_min, y_min, x_max, y_max] - ELLIPSE: [cx, cy, rx, ry] - POLYGON: [(x1, y1), (x2, y2), ...] (>= 3 points)

  • label (str) – Annotation label (e.g. "mass", "calcification").

  • metadata (dict | None) – Optional extra info. Defaults to {}.

Raises:

ValueError – If coordinates do not match the shape contract.

get_bounding_box()[source]

Return the axis-aligned bounding box enclosing the annotation.

Returns:

[x_min, y_min, x_max, y_max] in pixel coordinates.

Raises:

ValueError – For unsupported geometry types.

Return type:

List[int]

get_roi(padding=0, roi_type='bbox', image_shape=None)[source]

Return a region of interest around the annotation.

Computes the bounding box, applies padding, optionally clamps to image bounds, and returns the result in the requested shape.

Parameters:
  • padding (int) – Extra pixels added on each side of the bounding box.

  • roi_type (str) – Output shape format. "bbox" or "rectangle" returns {"type": ..., "coordinates": [x_min, y_min, x_max, y_max]}. "ellipse" returns {"type": "ellipse", "coordinates": {"center": (cx, cy), "radii": (rx, ry)}}.

  • image_shape (Tuple[int, int] | None) – (height, width) used to clamp coordinates so the ROI stays within image bounds. None means no clamping.

Returns:

dict with keys "type" (str) and "coordinates".

Raises:

ValueError – If roi_type is not "bbox", "rectangle", or "ellipse".

Return type:

dict

copy()[source]

Return an independent deep copy of this annotation.

Return type:

Annotation

to_dict()[source]

Serialize the annotation to a JSON-compatible dictionary.

The output includes computed fields (center, bounding_box) so that consumers do not need to recompute them. Polygon coordinates are converted from tuples to nested lists for JSON compatibility.

Returns:

dict with keys shape, coordinates, label, center, bounding_box, and metadata.

Return type:

dict

classmethod from_dict(data)[source]

Deserialize an annotation from a dictionary.

Inverse of to_dict(). The center and bounding_box fields in data are ignored (they are recomputed from coordinates).

Parameters:

data (dict) – Dictionary with at least shape, coordinates, and label keys. metadata is optional (defaults to {}).

Returns:

A new Annotation instance.

Return type:

Annotation

GeometryType

class medical_image.data.annotation.GeometryType[source]

Bases: Enum

Supported geometric shapes for annotations.

Each member defines the coordinate format expected by Annotation:

  • RECTANGLE – [x_min, y_min, x_max, y_max]

  • ELLIPSE – [cx, cy, rx, ry] (center + radii)

  • POLYGON – [(x1, y1), (x2, y2), ...] (>= 3 vertices)

  • BOUNDING_BOX – backward-compatible alias for RECTANGLE

RECTANGLE = 1
ELLIPSE = 2
POLYGON = 3
BOUNDING_BOX = 1