Skip to primary content
Technical Reference Architecture

Multi-Modal Vision OCR & LangGraph Clinical Extraction

Reviewed by Umar Abbas • Founder & Principal AI Architect

This technical reference architecture details the multi-modal vision OCR pipeline, LangGraph state machine orchestration, and post-mortem table segmentation fix for complex clinical document parsing. Engineered with AWS Bedrock Zero Data Retention endpoints and Detectron2 vision transformers, the blueprint evaluates layout parsing across multi-column medical charts.

Architecture PatternDetectron2 Vision OCR + LangGraph State Machine
Primary Constraint SolvedMulti-column table parsing without column bleeding
StackLangGraph, Detectron2, AWS Bedrock, Python
Data BasisPublic medical document datasets and synthetic chart records
1. Executive Summary & Build Context

Clinical Unstructured Data Bottlenecks

Architecture Note: This reference architecture documents a clinical extraction system engineered by Esaholic to parse multi-modal medical records without exposing Protected Health Information (PHI). Healthcare workflows process millions of unstructured clinical documents annually, including scanned physician notes, multi-column lab diagnostic charts, complex dosage tables, and faxed pathology reports.

Traditional Optical Character Recognition (OCR) engines lack spatial understanding of multi-column medical charts, leading to dosage misalignment and high manual review overhead. By combining Detectron2 layout transformers with LangGraph state machines and AWS Bedrock Zero Data Retention (ZDR) APIs, our team created a robust clinical extraction pipeline.

2. Problem & Baseline Bottlenecks

The Manual Medical Abstraction Crisis

Medical abstraction requires extracting discrete diagnostic codes and lab numbers from heterogeneous document layouts into electronic health record (EHR) systems. Manual data entry introduces error rates and creates operational backlogs that delay critical workflows.

Learn how our Document Processing Automation Solution Blueprint handles enterprise multi-page PDF ingestion with Pydantic validation.

Baseline Engineering Constraints
  • Layout Misalignment: Multi-column text bleeding across table boundaries in OCR output.
  • Low-Resolution Artifacts: Faxed pathology scans confusing gridlines with negative numerical signs.
  • Schema Non-Compliance: Raw LLM outputs failing strict FHIR JSON schema validations.
  • PHI Isolation: Requirement for edge-based redaction prior to cloud tokenization.
3. Architectural Solution

Detectron2 Vision Layout + LangGraph State Machine

The pipeline orchestrates multi-modal vision parsing with LangGraph state recovery, passing scrubbed token inputs into AWS Bedrock ZDR endpoints under an active Business Associate Agreement (BAA).

Multi-Modal Clinical Extraction Pipeline Architecture

Interactive Flow Diagram
Multi-Modal Clinical Extraction Pipeline Architecture Data flow across edge PHI masking, Detectron2 layout segmentation, LangGraph state validation, and FHIR JSON export. 1. Local PHI Scrubbing Edge SpaCy Engine 2. Vision OCR Layout Detectron2 Transformer 3. LangGraph Extraction AWS Bedrock ZDR 4. Validation & HITL Pydantic Guard 5. FHIR EHR Sync Zero Data Retention
Stage 1: 1. Local PHI Scrubbing Edge redaction

Identifies and redacts names, SSNs, and DOBs before cloud API transmission.

Data flow across edge PHI masking, Detectron2 layout segmentation, LangGraph state validation, and FHIR JSON export.
Text alternative for screen readers & search engines
Step Stage Name Function & Detail Metrics / SLA
1 1. Local PHI Scrubbing Identifies and redacts names, SSNs, and DOBs before cloud API transmission. Edge redaction
2 2. Vision OCR Layout Segments document layout into bounding-box region coordinates for text and tables. Layout parsing
3 3. LangGraph Extraction Extracts structured FHIR fields into Pydantic models with type enforcement. Structured schema
4 4. Validation & HITL Checks field confidence scores; routes low-confidence items to human review queues. Quality gate
5 5. FHIR EHR Sync Posts FHIR DiagnosticReport JSON to EHR and flushes volatile container RAM. Export complete
4. Technical Implementation

LangGraph State Machine Graph Code

Below is the LangGraph state graph enforcing Pydantic validation and routing low-confidence lab extractions to Human-in-the-Loop (HITL) interrupt nodes.

python / clinical_langgraph_pipeline.pyLangGraph State Graph & Pydantic Validator
from typing import TypedDict, List, Optional
from pydantic import BaseModel, Field
from langgraph.graph import StateGraph, END

class LabResultItem(BaseModel):
    test_name: str
    numeric_value: float
    unit: str
    reference_range: str
    confidence: float = Field(ge=0.0, le=1.0)

class ClinicalChartState(TypedDict):
    document_id: str
    raw_ocr_blocks: List[dict]
    parsed_labs: List[LabResultItem]
    needs_human_review: bool
    fhir_bundle: Optional[dict]

def extract_clinical_entities_node(state: ClinicalChartState) -> ClinicalChartState:
    # Simulating vision transformer structured extraction
    extracted_items = []
    low_confidence_flag = False
    
    for block in state["raw_ocr_blocks"]:
        item = LabResultItem(
            test_name=block.get("text_name", "Hemoglobin A1c"),
            numeric_value=float(block.get("val", 6.8)),
            unit="%",
            reference_range="4.0-5.6",
            confidence=float(block.get("score", 0.96))
        )
        if item.confidence < 0.92:
            low_confidence_flag = True
        extracted_items.append(item)
        
    state["parsed_labs"] = extracted_items
    state["needs_human_review"] = low_confidence_flag
    return state

def route_after_extraction(state: ClinicalChartState) -> str:
    if state["needs_human_review"]:
        return "hitl_interrupt_node"
    return "build_fhir_bundle_node"

# Constructing LangGraph State Graph
workflow = StateGraph(ClinicalChartState)
workflow.add_node("extract_entities", extract_clinical_entities_node)
workflow.add_conditional_edges(
    "extract_entities",
    route_after_extraction,
    {
        "hitl_interrupt_node": "hitl_interrupt_node",
        "build_fhir_bundle_node": "build_fhir_bundle_node"
    }
)
5. Technical Post-Mortem

What Went Wrong and How We Fixed It

Complex multi-modal document extraction across diverse scanned medical charts reveals edge-case failures during load testing. Here is our post-mortem analysis and resolution.

What Went Wrong: Scanned Gridline Confusion

During testing across low-resolution 150 DPI faxed pathology scans, standard vision OCR models misinterpreted horizontal table gridlines as negative minus signs. This corrupted numerical lab values by recording positive blood chemistry readings as negative numbers.

How We Fixed It: Bounding-Box Layout Anchors

We re-trained our Detectron2 vision model with explicitly labelled 2D bounding-box table coordinate masks and added a secondary Pydantic sanity-check guardrail node in LangGraph. This eliminated gridline token bleeding and restored numerical integrity.

6. Technical Evaluation Matrix

Standard OCR vs. Multi-Modal LangGraph Architecture

Engineering comparison evaluating the architectural trade-offs between standard single-pass OCR tools and the structured multi-modal pipeline.

Evaluation ParameterStandard Linear OCRMulti-Modal LangGraph ArchitectureArchitectural Benefit
Layout AwarenessSingle-stream linear text dump2D bounding-box spatial segmentationPreserves table column alignments
Schema ValidationUnstructured regex post-processingEnforced Pydantic type validationGuarantees valid FHIR output structure
Low-Confidence HandlingSilent failures and misreadsHITL interrupt routing nodesIsolates ambiguous charts for review
Data PrivacyDirect transmission of raw chartsEdge-based local PHI scrubbing with ZDRPrevents persistence of sensitive records

Note: Processing latency and throughput figures represent internal benchmarks conducted on synthetic medical charts in a local evaluation environment, not client production results.

7. Engineering Takeaways

Key Architectural Lessons

1. Bounding-Box Layout Pre-Segmentation

Segmenting unstructured PDFs into 2D layout blocks prior to sending text to LLMs prevents column bleeding and numerical table corruption.

2. Local Edge PHI Scrubbing

Redacting patient identifiers at the local edge network before invoking cloud AI endpoints guarantees strict data isolation even during API vendor outages.

3. LangGraph HITL Thresholds

Setting strict Pydantic confidence thresholds allows standard charts to process automatically while isolating low-confidence edge cases for human review.

9. Technical Blueprint FAQ

Frequently Asked Questions

How does the architecture ensure privacy during multi-modal vision OCR processing?↓

All incoming document pages pass through an edge-based local SpaCy/Presidio NER container that redacts patient identifiers prior to cloud vision transformer tokenization.

What caused the initial table extraction failure on scanned handwritten clinical notes?↓

Standard vision models missed non-standard table gridlines in low-resolution scans, causing misaligned lab values. Resolved by training a custom Detectron2 layout parser with bounding-box coordinate anchors.

What is the throughput capability of the multi-modal LangGraph pipeline?↓

The containerized cluster evaluates multi-page batches in parallel with average end-to-end extraction latency under 2 seconds per chart in local test environments.

How are ambiguous clinical abbreviations or illegible doctor notes handled?↓

When confidence scores fall below threshold limits, LangGraph triggers a Human-in-the-Loop interrupt node, routing the specific segment to medical verification queues.

Benchmark Your Document Extraction Architecture

Schedule a technical deep dive with Founder & Principal AI Architect Umar Abbas to review multi-modal OCR pipelines and LangGraph state machines under NDA.

Explore Agentic AI Services →