Multi-Modal Vision OCR & LangGraph Clinical Extraction
Reviewed by Umar Abbas • Founder & Principal AI Architect
This technical reference architecture details the multi-modal vision OCR pipeline, LangGraph state machine orchestration, and post-mortem table segmentation fix for complex clinical document parsing. Engineered with AWS Bedrock Zero Data Retention endpoints and Detectron2 vision transformers, the blueprint evaluates layout parsing across multi-column medical charts.
Clinical Unstructured Data Bottlenecks
Architecture Note: This reference architecture documents a clinical extraction system engineered by Esaholic to parse multi-modal medical records without exposing Protected Health Information (PHI). Healthcare workflows process millions of unstructured clinical documents annually, including scanned physician notes, multi-column lab diagnostic charts, complex dosage tables, and faxed pathology reports.
Traditional Optical Character Recognition (OCR) engines lack spatial understanding of multi-column medical charts, leading to dosage misalignment and high manual review overhead. By combining Detectron2 layout transformers with LangGraph state machines and AWS Bedrock Zero Data Retention (ZDR) APIs, our team created a robust clinical extraction pipeline.
The Manual Medical Abstraction Crisis
Medical abstraction requires extracting discrete diagnostic codes and lab numbers from heterogeneous document layouts into electronic health record (EHR) systems. Manual data entry introduces error rates and creates operational backlogs that delay critical workflows.
Learn how our Document Processing Automation Solution Blueprint handles enterprise multi-page PDF ingestion with Pydantic validation.
- Layout Misalignment: Multi-column text bleeding across table boundaries in OCR output.
- Low-Resolution Artifacts: Faxed pathology scans confusing gridlines with negative numerical signs.
- Schema Non-Compliance: Raw LLM outputs failing strict FHIR JSON schema validations.
- PHI Isolation: Requirement for edge-based redaction prior to cloud tokenization.
Detectron2 Vision Layout + LangGraph State Machine
The pipeline orchestrates multi-modal vision parsing with LangGraph state recovery, passing scrubbed token inputs into AWS Bedrock ZDR endpoints under an active Business Associate Agreement (BAA).
Multi-Modal Clinical Extraction Pipeline Architecture
Interactive Flow DiagramIdentifies and redacts names, SSNs, and DOBs before cloud API transmission.
Text alternative for screen readers & search engines
| Step | Stage Name | Function & Detail | Metrics / SLA |
|---|---|---|---|
| 1 | 1. Local PHI Scrubbing | Identifies and redacts names, SSNs, and DOBs before cloud API transmission. | Edge redaction |
| 2 | 2. Vision OCR Layout | Segments document layout into bounding-box region coordinates for text and tables. | Layout parsing |
| 3 | 3. LangGraph Extraction | Extracts structured FHIR fields into Pydantic models with type enforcement. | Structured schema |
| 4 | 4. Validation & HITL | Checks field confidence scores; routes low-confidence items to human review queues. | Quality gate |
| 5 | 5. FHIR EHR Sync | Posts FHIR DiagnosticReport JSON to EHR and flushes volatile container RAM. | Export complete |
LangGraph State Machine Graph Code
Below is the LangGraph state graph enforcing Pydantic validation and routing low-confidence lab extractions to Human-in-the-Loop (HITL) interrupt nodes.
from typing import TypedDict, List, Optional
from pydantic import BaseModel, Field
from langgraph.graph import StateGraph, END
class LabResultItem(BaseModel):
test_name: str
numeric_value: float
unit: str
reference_range: str
confidence: float = Field(ge=0.0, le=1.0)
class ClinicalChartState(TypedDict):
document_id: str
raw_ocr_blocks: List[dict]
parsed_labs: List[LabResultItem]
needs_human_review: bool
fhir_bundle: Optional[dict]
def extract_clinical_entities_node(state: ClinicalChartState) -> ClinicalChartState:
# Simulating vision transformer structured extraction
extracted_items = []
low_confidence_flag = False
for block in state["raw_ocr_blocks"]:
item = LabResultItem(
test_name=block.get("text_name", "Hemoglobin A1c"),
numeric_value=float(block.get("val", 6.8)),
unit="%",
reference_range="4.0-5.6",
confidence=float(block.get("score", 0.96))
)
if item.confidence < 0.92:
low_confidence_flag = True
extracted_items.append(item)
state["parsed_labs"] = extracted_items
state["needs_human_review"] = low_confidence_flag
return state
def route_after_extraction(state: ClinicalChartState) -> str:
if state["needs_human_review"]:
return "hitl_interrupt_node"
return "build_fhir_bundle_node"
# Constructing LangGraph State Graph
workflow = StateGraph(ClinicalChartState)
workflow.add_node("extract_entities", extract_clinical_entities_node)
workflow.add_conditional_edges(
"extract_entities",
route_after_extraction,
{
"hitl_interrupt_node": "hitl_interrupt_node",
"build_fhir_bundle_node": "build_fhir_bundle_node"
}
)What Went Wrong and How We Fixed It
Complex multi-modal document extraction across diverse scanned medical charts reveals edge-case failures during load testing. Here is our post-mortem analysis and resolution.
During testing across low-resolution 150 DPI faxed pathology scans, standard vision OCR models misinterpreted horizontal table gridlines as negative minus signs. This corrupted numerical lab values by recording positive blood chemistry readings as negative numbers.
We re-trained our Detectron2 vision model with explicitly labelled 2D bounding-box table coordinate masks and added a secondary Pydantic sanity-check guardrail node in LangGraph. This eliminated gridline token bleeding and restored numerical integrity.
Standard OCR vs. Multi-Modal LangGraph Architecture
Engineering comparison evaluating the architectural trade-offs between standard single-pass OCR tools and the structured multi-modal pipeline.
| Evaluation Parameter | Standard Linear OCR | Multi-Modal LangGraph Architecture | Architectural Benefit |
|---|---|---|---|
| Layout Awareness | Single-stream linear text dump | 2D bounding-box spatial segmentation | Preserves table column alignments |
| Schema Validation | Unstructured regex post-processing | Enforced Pydantic type validation | Guarantees valid FHIR output structure |
| Low-Confidence Handling | Silent failures and misreads | HITL interrupt routing nodes | Isolates ambiguous charts for review |
| Data Privacy | Direct transmission of raw charts | Edge-based local PHI scrubbing with ZDR | Prevents persistence of sensitive records |
Note: Processing latency and throughput figures represent internal benchmarks conducted on synthetic medical charts in a local evaluation environment, not client production results.
Key Architectural Lessons
Segmenting unstructured PDFs into 2D layout blocks prior to sending text to LLMs prevents column bleeding and numerical table corruption.
Redacting patient identifiers at the local edge network before invoking cloud AI endpoints guarantees strict data isolation even during API vendor outages.
Setting strict Pydantic confidence thresholds allows standard charts to process automatically while isolating low-confidence edge cases for human review.
Technologies & Services Used in This Build
Frequently Asked Questions
How does the architecture ensure privacy during multi-modal vision OCR processing?↓
All incoming document pages pass through an edge-based local SpaCy/Presidio NER container that redacts patient identifiers prior to cloud vision transformer tokenization.
What caused the initial table extraction failure on scanned handwritten clinical notes?↓
Standard vision models missed non-standard table gridlines in low-resolution scans, causing misaligned lab values. Resolved by training a custom Detectron2 layout parser with bounding-box coordinate anchors.
What is the throughput capability of the multi-modal LangGraph pipeline?↓
The containerized cluster evaluates multi-page batches in parallel with average end-to-end extraction latency under 2 seconds per chart in local test environments.
How are ambiguous clinical abbreviations or illegible doctor notes handled?↓
When confidence scores fall below threshold limits, LangGraph triggers a Human-in-the-Loop interrupt node, routing the specific segment to medical verification queues.
Benchmark Your Document Extraction Architecture
Schedule a technical deep dive with Founder & Principal AI Architect Umar Abbas to review multi-modal OCR pipelines and LangGraph state machines under NDA.
Explore Agentic AI Services →