What it does: Vision Builder scans your iPhone photo library, cuts every object out with SAM 2.1, groups look-alikes, lets you name each group once, and exports a labeled object-detection dataset (COCO, YOLO or CSV), all on the phone.
Who it is for: anyone who wants a labeled dataset of their own things, for example to train a detector for a robot, without uploading photos anywhere. You need an iPhone and Xcode to build it; see Requirements.
Status: work in progress. Some parts are rough and some are not wired up yet. The status table below says which.
Proof: real results on a 4,959-photo camera roll, the Swift source for every stage in this repo, and the model conversion scripts. There are no screenshots, demo video or App Store build yet.
Warning
This is an ongoing project. Things are half-built. Buttons sometimes lie. Empty states are awkward. The roadmap is longer than the README. We're shipping in the open because the core idea is too good to wait for "done."
If you find rough edges — that's the point. Tell us. Or fix it. PRs welcome.
|
🚀 Start here |
🔬 Under the hood |
🧑💻 Get going |
Imagine you just bought a robot. Or you're building one. It needs to recognize your stuff — your kitchen, your tools, your kid's toys, the specific objects in your house. Not generic ImageNet "person / car / dog." Your specific stuff.
The classical pipeline goes something like this:
- 1. Photograph thousands of objects from every angle
- 2. Upload to Roboflow / V7 / Scale AI
- 3. Pay someone to draw bounding boxes for weeks
- 4. Wait
- 5. Pay more
- 6. Train a model
- 7. Deploy+ Vision Builder collapses steps 1–5 into:
+ "Open the app. Walk around your house. Done."When this thing is "done" (it never will be — software is forever), here's what it'll do:
| Step | What happens |
|---|---|
| 📸 Capture | Take photos, or aim at your existing library — your phone already has thousands of pictures of your life |
| ✂️ Auto-segment | SAM 2 (or SAM 3, once a CoreML export exists) cuts every object out of every photo, automatically |
| 🧠 Auto-cluster | MobileCLIP 2 fingerprints each cut-out and groups them by visual similarity — "hey, here are 14 photos of what looks like the same coffee mug" |
| 🏷️ Label once | You tap one cluster, type "coffee mug," and all 14 instances inherit that label. Big-batch instead of one-photo-at-a-time grinding |
| 🤖 Export | Spit out a clean labeled dataset in COCO JSON, YOLO TXT, or CSV. Drop it straight into your model training pipeline |
| 🔒 Stay private | Nothing leaves the device. No cloud. No upload. No "free tier." Your photos, your data, your dataset |
The end state is a tool you walk around your space with for an afternoon and walk away with a personalized object-detection dataset ready to train a model that actually knows your world.
Not a benchmark. 4,959 photos off one real iPhone camera roll, three dogs the owner then identified by name.
|
"which dog is this"
✅ All three kept apart at cosine distance 0.75. |
"is this a dog"
|
Important
This is why the early clustering never worked, and it was never a tuning problem. CLIP is trained to tell a chair from a lamp — not your chair from mine. Identity needs an identity embedder. 🎯
Note
These groups were computed on a Mac with the same DINOv2-small model the app uses, then loaded into the app's Name tab (SeededInboxView.swift). The on-phone service (DINOv2Service.swift) exists but is not yet wired into the phone's own scan, which still clusters MobileCLIP embeddings with DBSCAN. The Mac-side grouping script is not in this repo.
| ❌ Limit | 🔍 What we measured |
|---|---|
| Identical mass-produced things | 28 different shipping labels over 886 days → one tight cluster. For inventory, identity isn't in the pixels at all. |
| Same-breed poultry | 🦆 Ducks separate from 🐔 chickens cleanly. Individual hens? No — and nothing fixes that. |
| People, full-body | Clusters by clothing and event, not person. That half needs a face embedder. |
Designed and written by Matt Macosko. The models are upstream (Meta SAM 2.1 and DINOv2, Apple MobileCLIP 2, Ultralytics YOLO 26 / YOLOE); the app and glue around them is mine:
- Library scan + clustering:
PhotoLibraryIndexer.swift,ObjectRecognitionEngine.swift(DBSCAN on MobileCLIP embeddings, pending clusters, label-once) - SAM 2.1 on CoreML:
SAM2CoreMLProcessor.swift,SAM2DetectionManager.swift - MobileCLIP 2 embeddings + tokenizer:
MobileCLIPService.swift,CLIPTokenizer.swift,scripts/convert_mobileclip2.py - DINOv2 identity + the dog test above:
DINOv2Service.swift,scripts/convert_dinov2.py - YOLOE 4,585-class detector with YOLO 26 / v8 fallback:
YOLOObjectDetector.swift,scripts/convert_yoloe.py,scripts/convert_yolo26.py - Naming flow:
MainTabView.swift(Name tab),SeededInboxView.swift,MorningInboxView.swift - Live camera recognition:
LiveRecognitionView.swift - Bird's-eye mosaic + concept search:
BirdsEyeView.swift,ConceptSearchService.swift - COCO / YOLO / CSV export:
ExportManager.swift,ExportOptionsView.swift - Model conversion pipeline (Python 3.12 venv, coremltools patch, fp16 NaN checks):
scripts/convert_models.sh
Because every "AI for robotics" tutorial assumes you have a labeling team.
Most people don't. They have a phone and 15 minutes between things.
Because Roboflow is great, but uploading 4,000 photos of your living room to a cloud service feels insane.
Because the iPhone Neural Engine is faster than most cloud GPUs were five years ago, and it's just sitting in your pocket.
Because the future of robots learning to operate in your environment shouldn't require a SaaS subscription.
flowchart LR
A[📸 Photo Library] -->|scan| B[✂️ SAM 2.1<br/>segment objects]
B --> C[🧠 MobileCLIP 2<br/><i>what</i> is it]
B -.->|Mac-side today| H[🆔 DINOv2<br/><i>which one</i> is it]
C --> D[📊 DBSCAN cluster]
H -.-> E
D --> E[🏷️ Name tab<br/>name it once]
E --> F[🤖 Export<br/>COCO / YOLO / CSV]
F --> G[🦾 Train your robot]
style A fill:#FF9500,stroke:#fff,color:#fff
style B fill:#FF2D55,stroke:#fff,color:#fff
style C fill:#AF52DE,stroke:#fff,color:#fff
style H fill:#FF375F,stroke:#fff,color:#fff
style D fill:#5856D6,stroke:#fff,color:#fff
style E fill:#007AFF,stroke:#fff,color:#fff
style F fill:#34C759,stroke:#fff,color:#fff
style G fill:#00C7BE,stroke:#fff,color:#fff
Note
Status legend — ✅ working · 🟡 functional but rough · 💤 scaffolded, dormant · ❌ not yet
| Component | Status | Notes |
|---|---|---|
| SAM 2.1 segmentation | ✅ | Solid, runs on Neural Engine |
| MobileCLIP 2 embeddings | ✅ | fp32 (the fp16 tower NaNs — see convert script), center-crop preprocessing |
| YOLOE detection | ✅ | 4,585 classes, prompt-free open-vocab, decode baked into the CoreML graph (YOLO 26 / v8 fallback) |
| Live recognition tab | ✅ | Point the camera — objects you've labeled get named on screen, on-device |
| Bird's-eye dataset mosaic | ✅ | Whole dataset on one zoomable wall (My Things → ⋯) |
| DBSCAN clustering | ✅ | Euclidean on unit vectors, eps measured against real embeddings (not vibes) |
| DINOv2 identity grouping | 🟡 | Measured on a Mac (numbers above); on-phone service written but not wired into the scan yet |
| Photo library scan | ✅ | Working but slow on big libraries |
| Name tab cluster review | 🟡 | Functional, UX still rough |
| Manual labeling flow | 🟡 | Has back/next now, editor is busy |
| Concept search | 🟡 | "Find all my cups" — works, sometimes underwhelming |
| COCO / YOLO / CSV export | ✅ | Ready for any standard pipeline, real zips, share sheet |
| Foundation Models smart-naming | 💤 | Scaffolded — needs A17 Pro+ device |
| SAM 3 text-prompted segmentation | 💤 | Skeleton ready, waiting on an upstream EfficientSAM3 CoreML export |
| iCloud sync | ❌ | On the list |
| Apple Watch quick-label | ❌ | On the list |
| Siri Shortcuts | ❌ | On the list |
|
Framework layer
|
ML layer
|
Apple Intelligence
|
Important
- iOS 26.0+ (Foundation Models APIs + iOS 26 SwiftUI bits)
- iPhone with A14 chip or later for Neural Engine acceleration
- Apple Intelligence device (iPhone 15 Pro+) only for the on-device LLM cluster-naming — everything else runs on older iPhones
- Xcode 26+ to build (the project targets the iOS 26 SDK)
- A Mac with Apple Silicon + Python 3.10 to 3.12 only if you want to regenerate the bigger models
git clone https://github.com/nicedreamzapp/VisionBuilder.git
cd VisionBuilder
# Optional: generate the upgraded mlpackages (MobileCLIP 2 + YOLO 26)
# Takes ~5 min on Apple Silicon
bash scripts/convert_models.sh all
# Optional extras, not included in "all":
bash scripts/convert_models.sh yoloe # 4,585-class YOLOE detector
source .venv-models/bin/activate # created by convert_models.sh
python3 scripts/convert_dinov2.py # writes dinov2_small_fp16.mlpackage (and an fp32 variant); add it to the target
# Open in Xcode and build to a real device
# (Simulator works but it's slow — no Neural Engine)
open "Vision Builder.xcodeproj"Tip
The conversion script auto-creates a Python 3.12 venv (coremltools 8 hates Python 3.14), patches a known coremltools bug for newer numpy, and produces:
mobileclip2_s0_image.mlpackage(22 MB)mobileclip2_s0_text.mlpackage(121 MB)yolo26n.mlpackage(4.8 MB)
These are gitignored — too big for GitHub's 100MB file limit, but reproducibly regenerated by the script. Drag them into Xcode and add them to the Vision Builder target.
If you skip conversion, the app still builds and runs on the models already committed here: SAM 2.1, MobileCLIP S0 (v1) and YOLOv8n Open Images. It picks YOLOE, then YOLO 26, then YOLOv8 based on what is bundled.
1. Grant photo library access
2. My Things tab → ⋯ → "Scan Photo Library"
3. Wait — you'll see the object count climb as it scans
4. Name tab → label each cluster (one tap per cluster of N similar objects)
5. My Things tab → ⋯ → "Export All" → COCO/YOLO/CSV
6. Feed the dataset into your training pipeline
7. Train. Deploy. Profit. (lol, jk, you'll iterate forever)
Note
The app also has developer test hooks for a Mac (PhotoExportView.swift, RemoteCaptureView.swift, PhotoCleanupView.swift). They only run when a Mac has dropped a request file into the app's Documents folder (copied with devicectl). Normal use never sends anything off the phone.
Note
Rough priority order. Subject to change every time we open the app and trip over a bug.
- Make the Inbox flow buttery — it's the heart of the app
- SAM 3 integration once an EfficientSAM3 CoreML export ships upstream
- Wire DINOv2 into the phone's own scan so identity grouping no longer needs the Mac
- iCloud sync so you can label across iPhone + iPad
- Apple Watch companion for "label this object" in the wild
- Siri Shortcuts ("scan my latest 100 photos")
- On-device YOLO finetuning — start with COCO, finetune to your dataset right on the phone
- Multi-project support — one dataset for kitchen, another for shop, another for the robot's path
- Smarter cluster auto-naming with Foundation Models when device supports
- Collaborative datasets without cloud — maybe AirDrop the
.mlpackage? - Robot-specific export presets — URDF-aware? object pose? still scoping
Genuinely, please.
| If you... | Do this |
|---|---|
| Found a 🐛 bug | Open an issue — even rough notes help |
| Have a 💡 idea | Start a discussion |
| Want to 🔧 code | PRs welcome — file structure is mostly self-explanatory, see CLAUDE.md |
I'm learning as I go. If you're a robotics person with opinions about training-data formats, an iOS dev who's done CoreML in anger, or just someone who's tried to label 400 photos of their dog and hated it, you'd add value here.
| Who | What for |
|---|---|
| 🦾 Meta AI | SAM 2.1 and DINOv2 — segmenting anything is wild |
| 🍎 Apple ML Research | MobileCLIP 2 — vision-language tiny enough for a phone |
| 🚀 Ultralytics | YOLO 26 + YOLOE, and actually shipping CoreML exports |
| 🤖 Roboflow / Scale AI / V7 | Showing what good labeling tools look like, even if we do it on-device instead |
MIT. Do what you want with it. If you build something cool, tell us — we want to see the robot.
Built with ❤️ + 🤖 + a healthy distrust of cloud services.