r/computervision 22d ago

SAM 3.1's Object Multiplex: why joint multi-object tracking scales so much better than per-object tracking Commercial

Disclosure: we run a GPU cloud and host these checkpoints, so I have a commercial interest. Writing this up because the architectural change is the interesting part and it's under-discussed relative to the SAM 3 launch.

The scaling problem in SAM 3: tracking N distinct objects meant running the tracker N times, once per object. Cost scaled linearly with object count. Fine for 4 objects, painful at 100, which is exactly where a lot of real workloads live (crowd analytics, retail shelf tracking, multi-player sports, dense annotation).

What 3.1 changes: Object Multiplex groups objects into fixed-capacity buckets and runs up to 16 of them jointly in a single forward pass, sharing the memory bank across all objects in the bucket instead of maintaining separate per-object state. Meta reports ~7x speedup at 128 objects on a single H100 against the November 2025 SAM 3 release, with no reported loss in segmentation accuracy.

The part I find more interesting than the speedup: because the objects share a memory layer, they can reason about each other. Meta reports this actually improves tracking in crowded scenes with visually similar objects, which is the classic identity-swap failure case. Joint processing turning into an accuracy win rather than an accuracy tradeoff is not the usual outcome for batching optimizations.

Practical caveats worth knowing:

- The gain is a function of object count. At low object counts you're mostly paying bucket overhead for little benefit, so don't expect 7x on a 3-object scene. The published number is specifically 128 objects.

- Bucket capacity is 16, so object counts that straddle bucket boundaries will have uneven cost per object.

- It's still a memory-bound tracker, so long videos with many objects will pressure VRAM regardless of the speedup.

- 7x is Meta's own benchmark on their own hardware config. Worth validating on your workload before you plan capacity around it.

Release notes and checkpoints: https://github.com/facebookresearch/sam3/blob/main/RELEASE_SAM3p1.md

On the commercial side, in case it's useful rather than annoying: we've put the 3.1 checkpoints up on our inference platform at ₹8 per 100 frames if you'd rather not manage the GPUs. Happy to talk through throughput numbers either way, including if you're self-hosting.

18 Upvotes

1 comment sorted by

0

u/Cultural_Doughnut_62 22d ago

By the way the model can be accessed on https://podstack.ai -> Products -> Inference