Segment Anything (SAM): Promptable, Class-Agnostic Segmentation at Scale
Reviewed by Umar Abbas • Founder & Principal AI Architect
Last reviewed: 14 August 2026
Segment Anything (SAM) is Meta's promptable segmentation foundation model. Given point, box, or mask prompts, it produces high-quality class-agnostic masks without task-specific training. Its heavy ViT image encoder runs once per image, then a lightweight decoder returns masks in milliseconds, enabling zero-shot, interactive segmentation.
What Segment Anything Solves in Production
Traditional segmentation pipelines require a labeled dataset and a training cycle for every new object class, which is slow and expensive when the taxonomy shifts or a new part appears on the line. SAM removes that per-class training requirement by producing high-quality masks from simple geometric prompts, so a detector box or a single click is enough to isolate an object. This turns segmentation into a composable step rather than a bespoke model. It is especially useful when object classes are open-ended or annotation budgets are tight. The tradeoff is that SAM gives you geometry, not identity, so it must be combined with a labeling source.
Inside the SAM Architecture
Anatomy ExplainerCore Component Component Parts:
Image Encoder
A Vision Transformer that maps the input image into a dense embedding computed once per image.
MAE-pretrained ViT (H, L, or B) producing a 256-channel, 64 by 64 embedding at 1024 pixel input. This is the dominant compute cost and is cached for reuse across prompts.
Text alternative for screen readers & search engines
- Part 1: Image Encoder - A Vision Transformer that maps the input image into a dense embedding computed once per image. [Tech: MAE-pretrained ViT (H, L, or B) producing a 256-channel, 64 by 64 embedding at 1024 pixel input. This is the dominant compute cost and is cached for reuse across prompts.]
- Part 2: Prompt Encoder - Converts sparse and dense prompts into embeddings the decoder can consume. [Tech: Points and boxes are encoded with positional encodings plus learned type embeddings. Mask prompts are downsampled and added to the image embedding. Text prompts were a research capability, not a shipped default.]
- Part 3: Mask Decoder - A lightweight transformer that fuses image and prompt embeddings into output masks. [Tech: Two-way cross attention between prompt tokens and image embedding, followed by an upsampling head. Runs in roughly 50 milliseconds, enabling interactive prompting after the encode step.]
- Part 4: Ambiguity Head - Resolves the many valid masks a single prompt can imply. [Tech: With multimask_output enabled the decoder emits three masks per prompt at part, object, and scene granularity, each with a predicted IoU confidence used for ranking or user selection.]
- Part 5: Automatic Mask Generator - A wrapper that segments an entire image without manual prompts. [Tech: SamAutomaticMaskGenerator samples a grid of foreground points, runs the decoder per point, then filters by predicted IoU and stability score and removes duplicates with non-maximum suppression.]
Architectural Strengths & Specific Production Limits
- Zero-shot generalization: SAM produces usable masks on unseen object types without any fine-tuning, which collapses the cost of adding new categories.
- Decoupled compute: Encoding once and prompting many times makes interactive click-to-segment tools responsive after the initial embedding is cached.
- Prompt flexibility: Points, boxes, and coarse masks can all drive segmentation, so SAM slots cleanly behind a detector or a human annotator.
- Permissive licensing: Apache 2.0 model weights and code allow commercial deployment without the restrictions attached to some research releases.
- No semantic labels: SAM returns geometry only and never names the object, so a separate classifier or detector is mandatory when class identity matters.
- Heavy default encoder: The ViT-H image encoder is too large for most real-time edge use, forcing distilled variants or server-side embedding precomputation.
- Domain gaps: Accuracy drops on imagery far from natural photos, such as low-contrast medical or dense satellite scenes, without a fine-tuned decoder.
- Fine boundary noise: Thin structures, transparency, and heavily occluded edges can produce ragged or leaking masks that need post-processing.
How We Deploy Segment Anything in Production
Our default pattern treats SAM as the mask engine behind a detector rather than a standalone system. A lightweight detector supplies box prompts, SAM converts them to precise instance masks, and the encoder runs on batched frames on a GPU worker with embeddings cached so downstream prompts stay cheap. For interactive tooling we export the decoder to ONNX and serve embeddings from an inference service, keeping the browser side responsive. We pin checkpoints and library versions, and we validate mask quality on held-out domain samples before promotion.
SAM Serving Pipeline
Interactive Flow DiagramFrames are resized to the 1024 pixel long side and normalized to match the encoder input distribution.
Text alternative for screen readers & search engines
| Step | Stage Name | Function & Detail | Metrics / SLA |
|---|---|---|---|
| 1 | 1. Ingest and Preprocess | Frames are resized to the 1024 pixel long side and normalized to match the encoder input distribution. | 1024px input |
| 2 | 2. Encode | The ViT encoder runs once per image on GPU and the resulting embedding is cached keyed by frame hash. | one encode per image |
| 3 | 3. Prompt | A detector supplies box prompts, or a user click supplies point prompts, which the prompt encoder embeds. | boxes or points |
| 4 | 4. Decode Masks | The mask decoder returns candidate masks with IoU scores, selecting the highest confidence mask per prompt. | approx 50ms per prompt |
| 5 | 5. Label and Post-process | Detector class labels are joined to masks, then morphology and NMS clean boundaries before serialization. | labels joined |
# pip install segment-anything==1.0 torch==2.1.0 opencv-python numpy
import numpy as np
import cv2
import torch
from segment_anything import sam_model_registry, SamPredictor
DEVICE = "cuda" if torch.cuda.is_available() else "cpu"
MODEL_TYPE = "vit_h"
CHECKPOINT = "sam_vit_h_4b8939.pth"
sam = sam_model_registry[MODEL_TYPE](checkpoint=CHECKPOINT)
sam.to(device=DEVICE)
predictor = SamPredictor(sam)
image = cv2.cvtColor(cv2.imread("frame.jpg"), cv2.COLOR_BGR2RGB)
predictor.set_image(image) # heavy encoder runs once, embedding cached
# a box prompt from an upstream detector, format x0 y0 x1 y1
box = np.array([220, 140, 640, 520])
masks, scores, logits = predictor.predict(
box=box,
multimask_output=False,
)
best_mask = masks[0]
print("mask pixels:", int(best_mask.sum()), "score:", float(scores[0]))Services Engineered with Segment Anything (SAM)
We help teams stand up and operate SAM-based segmentation in real deployments.
Segment Anything vs Alternative Segmentation Tools
How SAM compares to established supervised instance segmentation frameworks.
Segmentation Approach Matrix
Benchmark Matrix| Evaluation Metric | Segment Anything (SAM) | Detectron2 | YOLO Seg |
|---|---|---|---|
| Zero-shot generalization | Strong, no training needed Winner | Requires labeled training | Requires labeled training |
| Built-in class labels | None, class-agnostic | Full instance classes Winner | Full instance classes |
| Real-time inference | Heavy encoder | Moderate | Fast, edge capable Winner |
| Interactive promptability | Points and boxes Winner | Not interactive | Not interactive |
Text alternative for screen readers & search engines
- Zero-shot generalization: Segment Anything (SAM): Strong, no training needed vs Detectron2: Requires labeled training vs YOLO Seg: Requires labeled training (Winning option: Segment Anything (SAM)).
- Built-in class labels: Segment Anything (SAM): None, class-agnostic vs Detectron2: Full instance classes vs YOLO Seg: Full instance classes (Winning option: Detectron2).
- Real-time inference: Segment Anything (SAM): Heavy encoder vs Detectron2: Moderate vs YOLO Seg: Fast, edge capable (Winning option: YOLO Seg).
- Interactive promptability: Segment Anything (SAM): Points and boxes vs Detectron2: Not interactive vs YOLO Seg: Not interactive (Winning option: Segment Anything (SAM)).
Segment Anything (SAM) in a Reference Architecture
On a fintech document automation engagement we used promptable segmentation to isolate stamps, signatures, and table regions on scanned forms without training a class per document layout. Class-agnostic masks let the pipeline adapt to new form variants quickly, with a downstream classifier attaching meaning to each cropped region.
Read Reference Architecture →Frequently Asked Questions
What is Segment Anything (SAM)?↓
SAM is a segmentation foundation model released by Meta AI in April 2023. It takes prompts such as points or boxes and returns pixel-accurate masks for objects it has never been explicitly trained to recognize. It was trained on SA-1B, a dataset of 11 million images and over 1 billion masks.
Does SAM classify objects or name them?↓
No. SAM is class-agnostic. It outputs a mask for the region indicated by a prompt but does not assign a semantic label such as car or person. Teams typically pair SAM with a detector or a text encoder like CLIP when class names are required.
What is the difference between the ViT-H, ViT-L, and ViT-B checkpoints?↓
They are three image encoder sizes with a quality and speed tradeoff. ViT-H is the default 636 million parameter checkpoint with the strongest masks. ViT-B is the smallest and fastest, and ViT-L sits between them. All three share the same lightweight prompt encoder and mask decoder.
How fast is SAM at inference?↓
The cost is dominated by the image encoder, which typically takes a few hundred milliseconds to around a second on a modern GPU for ViT-H. After the embedding is computed once, each prompt runs through the mask decoder in roughly 50 milliseconds, which supports interactive editing loops.
What license does SAM use?↓
The Segment Anything model code and checkpoints are released under the Apache 2.0 license, which permits commercial use. The SA-1B training dataset has its own separate research-oriented terms, so review those before using the raw data rather than the model.
What is the difference between SAM and SAM 2?↓
SAM handles single images. SAM 2, released in 2024, extends the approach to video by adding a memory mechanism that propagates masks across frames for object tracking. SAM 2 also improves single-image speed while keeping the promptable interface.
Can SAM run in real time or on the edge?↓
The full ViT-H encoder is heavy for edge devices. Distilled variants such as MobileSAM, FastSAM, and EfficientSAM trade some quality for far smaller encoders. A common production pattern is exporting the decoder to ONNX and precomputing embeddings server side.
How does SAM handle prompt ambiguity?↓
A single point can refer to a part, an object, or a whole scene. SAM addresses this by setting multimask_output to true, which returns three candidate masks with confidence scores. Applications either pick the highest score or let a user choose the intended level.
How do you get masks for every object automatically?↓
Use SamAutomaticMaskGenerator, which samples a regular grid of point prompts across the image, runs the decoder for each, then filters and deduplicates the results with non-maximum suppression. It is thorough but heavier than a single prompted call.
Is SAM good for medical or satellite imagery out of the box?↓
Results vary. SAM generalizes well to natural images but underperforms on domains far from its training distribution, such as low-contrast medical scans or dense overhead imagery. Fine-tuned variants like MedSAM address this by adapting the decoder to the target domain.