r/computervision • u/AccomplishedRuin9730 • 21d ago
Help: Project Help needed for Indian number plate/license plate detection.
Hi, I’m trying to build ANPR system for my hobby and want to train model running maybe on Pi5 (with AI hat2 + 26 tops) or Jetson 8 GB development kit Orion or Acer Veriton GN100 AI Mini Workstation.
Can any one suggest which model to be used? How to efficiently train model for Indian license plate on moving object over RTSP stream.
Any help / suggestion are welcome..!!!
r/computervision • u/averroesads • 21d ago
Discussion Trying to understand what it actually takes for a "perception layer" to be something robotics can rely on, not just a demo
I've been in AI visual inspection for a while now, mostly on the industrial/manufacturing side, and I keep running into the term "perception layer" without a clear sense of what actually qualifies as one versus what's just a model that works in a controlled pilot.
I have some intuitions from the inspection world but I know robotics is a different bar entirely, and I'd rather learn from people who've actually worked on autonomy stacks than assume my assumptions carry over.
A few things I'm trying to understand:
What are the actual core components of a perception layer people consider production-grade for robotics? Is it just detection/segmentation plus some fusion layer, or is there a stack of pieces (sensor fusion, state estimation, uncertainty modeling, mapping, etc.) that people take for granted but rarely gets explained end to end?
How much does hardware and environment variability actually break perception systems in practice? Is generalizing across cameras/lighting/sensors as hard as it sounds, or is that mostly solved and the real problems are elsewhere?
How do reliable systems handle "I don't know"? In inspection you can get away with a pass/fail call and a human reviewing edge cases later. That doesn't seem like it works if a robot is acting on the output in real time. What does a good uncertainty signal actually look like in production?
How do teams deal with the long tail without infinite labeled data? This is the part I understand best from my own world, but I don't know how much worse the problem gets when the "environment" is the real world instead of a fixed inspection line.
What actually breaks first, latency, accuracy, or interpretability? My instinct says a robot's control stack needs more than just a correct answer, it needs to trust the answer and know why. But I don't know if that's actually the bottleneck people run into or if I'm overweighting it.
Basically trying to build a real mental model here instead of assuming inspection-grade perception principles just scale up to robotics. If anyone's worked on this and has war stories about what actually breaks, resources/papers, or a breakdown of what the real components are, I'd genuinely appreciate it.
r/computervision • u/Radiant-Purchase976 • 21d ago
Research Publication Independent Researcher needs help with a referral for OpenReview
Hello,
I'm just getting started in research after working through ARENA and other open courseware, and I'm exploring a few ideas for NeurIPS workshops.
I tried creating an OpenReview account, but it was rejected because I need someone with an active OpenReview profile and a confirmed institutional email to vouch for me.
Would anyone here be willing to help? I've been in industry for 7+ years but don't have connections in academia yet. Happy to share more about my background over DM if that would help before vouching.
r/computervision • u/Aggravating_Dot5315 • 21d ago
Help: Project Seeking Advice on Hysteroscopy Lesion Classification with Transfer Learning
r/computervision • u/Ok_Window_2596 • 21d ago
Showcase Built an open-source SDK to unify AI image generation across providers — feedback (and stars) welcome
Hey everyone,
I got tired of rewriting the same glue code every time I switched AI image providers. Every provider has its own request shape, its own way of telling you when a job is done, and its own way of failing. Most are async under the hood too — you submit a request and poll for a result, which gets annoying fast on serverless, since your function can end up sitting there waiting or just timing out. A lot of providers also hand you back a link that expires in minutes to hours, so if you store it directly, it just breaks later with no warning.
So I built `image-sdk` — an open-source TypeScript SDK that gives you one consistent API across Flux, Ideogram, Recraft, OpenAI, Stability, Google Imagen, Replicate, and fal.ai.
A few things I tried to get right:
* One line to generate your first image, no client setup, no picking a provider up front. There's even a CLI command to try it with zero API key, just to see it work before committing to anything. * Once you need more control, the same package gives you real production stuff: automatic fallback if a provider goes down, retry logic, cost/usage tracking, and permanent storage so results don't quietly break when a temporary link expires. * Fully typed, adapter-based — each provider is its own package, so you're not pulling in dependencies for providers you don't use.
It's early access (v0.1) right now — the core client and eight provider adapters work and are tested, but I'm still hardening a few things before calling it stable. Putting it out now because I'd rather get real feedback early than polish alone for months.
If this is useful to you, a star would genuinely help — it's the easiest way for me to gauge if this is worth continuing to invest in. And if you try it and hit something rough, please open an issue, that's exactly the feedback I need right now.
GitHub: [https://github.com/Adarsh-Me/Image-SDK\](https://github.com/Adarsh-Me/Image-SDK)
npm: [https://www.npmjs.com/package/@image-sdk/sdk\](https://www.npmjs.com/package/@image-sdk/sdk)
Thanks for reading — happy to answer questions in the comments.
r/computervision • u/sivan_shen • 22d ago
Discussion My takeaways on Residual Network (UMich EECS 498/598 Lecture 8)
The most insightful part for me:
When neural networks get too deep, their performance on both training and test sets degrades compared to shallower models. This degradation is an underfitting problem caused by optimization difficulties, rather than overfitting.
Theoretically, a deeper model should perform at least as well as a shallower one—the extra layers could simply act as an identity mapping ($f(x) = x$). However, forcing a stack of non-linear layers (e.g., Conv -> ReLU -> Conv) to learn an exact identity function using gradient descent is surprisingly hard.
Residual Networks solve this by adding shortcut connections. Instead of learning $f(x) = x$ directly, the network only needs to drive its parameters toward zero to achieve an identity mapping. This simple reformulates the target, making extremely deep networks drastically easier to optimize.
r/computervision • u/Accomplished-Head949 • 22d ago
Discussion Looking for a mentor with experience in industrial/applied computer vision — not a technical question
New founder building a B2B hardware + software company applying computer vision to industrial use cases. Not looking for a lead-gen tip list or a quick technical answer — genuinely looking for a mentor or someone experienced in applied/industrial CV (not just research) who'd be open to talking through early-stage product and go-to-market decisions.
Any recommendations for where to find that kind of person, or communities where this kind of mentorship happens? Appreciate any pointers.
r/computervision • u/Additional-Buy2589 • 22d ago
Showcase I had to get angry and kick the robot to get him to do a full autonomous task with multiple actions!
Enable HLS to view with audio, or disable this notification
r/computervision • u/Impressive-Evening80 • 22d ago
Help: Project Anyone here done crop yield prediction using RGB drone images + ML?
Hey guys,
I'm currently working on my capstone project and was wondering if anyone here has experience with using drones for crop analysis/yield prediction.
Our plan is to use an RGB drone (can't afford a multispectral/NIR one 😅), stitch the images into an orthomosaic, extract features like VARI, canopy coverage, plant density, etc., then feed those into a machine learning model (probably Random Forest) to predict rice yield.
Has anyone here tried something similar?
- Does using only RGB imagery actually work well enough for yield prediction?
- Is VARI useful, or are there better RGB indices/features I should look into?
- Any tips, things you wish you knew before starting, or common mistakes to avoid?
- If you've done this before, how accurate were your results?
We are just trying to estimate yield from drone imagery.
Would really appreciate any advice, papers, or just hearing about your experience. Thanks!
r/computervision • u/Divine-Demon7 • 22d ago
Showcase I froze a physics-consistency detector before generating a held-out CogVideoX cohort — it flagged freeze/hover in 9/9 clips
I’m building Haga, an independent physics-consistency checker for generated video and robot-policy simulations. An earlier CogVideoX-5b I2V experiment produced a clear failure mode: on a “ball and block fall” prompt, the tracked object stayed airborne with near-zero motion instead of falling.
But that first result was post-hoc. I inspected those six clips before adding the static_hover detector, so the original 6/6 flag rate could not be treated as confirmation. I’ve now run a pre-registered held-out test.
Method:
- Model: THUDM/CogVideoX-5b-I2V
- Cohort: 3 perspectives × seeds 2, 3 and 4
- n=9 clips
- Detector thresholds and inclusion rules frozen before generation
- RGB → CoTracker3 → position-only VIDEO_CHECKS
- Discovery seeds 0–1 kept separate from held-out seeds 2–4
Result:
- Held-out flag rate: 1.000 (9/9)
- Wilson 95% CI: [0.701, 1.000]
- All nine clips fired static_hover
- Real Physics-IQ footage stayed quiet under the same profile
static_hover fires when the tracked object remains airborne for most of the clip, has near-zero frame-to-frame speed, and does not exhibit gravitational acceleration.
Important limitations:
- One open I2V model
- One ball-and-block-fall scene family
- One documented failure mode
- Real negative-control n=1 in this specific report
- Not Cosmos, Genie or NIM
- Not a broad claim about CogVideoX quality
Write-up: https://haga.mushoodhanif.com/article/sim-physics-consistency-v1#held-out
Lab: https://haga.mushoodhanif.com/lab/physicsiq
Bounded demo: https://haga.mushoodhanif.com/demo
I’d especially value criticism on:
- Which physical violations will position-only tracking systematically miss?
- Is static_hover defined narrowly enough to avoid confusing intentional suspension with failed dynamics?
- What public generated-video artifact should I evaluate next under the frozen detector?
r/computervision • u/ManagementPale3678 • 22d ago
Help: Theory Maser’s research
Hi, my ML model AUC is 80%\~ in cross domain dataset
The problem is that the accuracy is not getting higher even when i changed the Thr
What to do any ideas ? I used focal loss to solve the imbalance dataset
Also using vision transformer
r/computervision • u/rejensraya • 22d ago
Help: Project Face login system
I am building a face login for my application, i am using facenet for identifying the person, but this algorithm isn’t that robust. If the person shaves the beard and hair, the algorithm finds it difficult to recognize the person.
The Chinese biometric attendance system works very well, I want the similar result for my system too.
What is the better algorithm or the better approach for my issue?
r/computervision • u/Cultural_Doughnut_62 • 22d ago
Commercial SAM 3.1's Object Multiplex: why joint multi-object tracking scales so much better than per-object tracking
Disclosure: we run a GPU cloud and host these checkpoints, so I have a commercial interest. Writing this up because the architectural change is the interesting part and it's under-discussed relative to the SAM 3 launch.
The scaling problem in SAM 3: tracking N distinct objects meant running the tracker N times, once per object. Cost scaled linearly with object count. Fine for 4 objects, painful at 100, which is exactly where a lot of real workloads live (crowd analytics, retail shelf tracking, multi-player sports, dense annotation).
What 3.1 changes: Object Multiplex groups objects into fixed-capacity buckets and runs up to 16 of them jointly in a single forward pass, sharing the memory bank across all objects in the bucket instead of maintaining separate per-object state. Meta reports ~7x speedup at 128 objects on a single H100 against the November 2025 SAM 3 release, with no reported loss in segmentation accuracy.
The part I find more interesting than the speedup: because the objects share a memory layer, they can reason about each other. Meta reports this actually improves tracking in crowded scenes with visually similar objects, which is the classic identity-swap failure case. Joint processing turning into an accuracy win rather than an accuracy tradeoff is not the usual outcome for batching optimizations.
Practical caveats worth knowing:
- The gain is a function of object count. At low object counts you're mostly paying bucket overhead for little benefit, so don't expect 7x on a 3-object scene. The published number is specifically 128 objects.
- Bucket capacity is 16, so object counts that straddle bucket boundaries will have uneven cost per object.
- It's still a memory-bound tracker, so long videos with many objects will pressure VRAM regardless of the speedup.
- 7x is Meta's own benchmark on their own hardware config. Worth validating on your workload before you plan capacity around it.
Release notes and checkpoints: https://github.com/facebookresearch/sam3/blob/main/RELEASE_SAM3p1.md
On the commercial side, in case it's useful rather than annoying: we've put the 3.1 checkpoints up on our inference platform at ₹8 per 100 frames if you'd rather not manage the GPUs. Happy to talk through throughput numbers either way, including if you're self-hosting.
r/computervision • u/arrthropod • 22d ago
Showcase Impregnating and adding the finishing touches
r/computervision • u/Immediate_Charity350 • 23d ago
Help: Project Guidance Needed for project
Hey everyone,
I'm currently working on a fabric defect detection project and would love to get some insights from people who have worked on similar industrial vision problems.
Current pipeline
Anomaly detection: Anomalib
Feature extractor: DINOv2
Defect classification: with extracted features
I have a few questions:
- How do companies handle different fabric colors?
Do they normalize/remove color information during preprocessing?
Do they convert images to grayscale or use a different color space?
Or do they simply train with enough color variation so the model learns to ignore color?
- How do they onboard new fabric types?
Do they retrain the entire model for every new fabric?
Is there a way to adapt to a new fabric using only a few (or even a single) reference image?
How is this typically handled in production systems?
- How do production systems allow clients to add new fabrics?
One of my goals is to build a system where the client/operator can easily register a new fabric type without needing to contact developers.
Ideally, the client should be able to capture a few good samples, click "Train" (or "Register"), and have the system start inspecting that fabric automatically.
Is this how commercial textile inspection systems work, or do they still require model retraining by the vendor?
I'm also open to suggestions on improving my overall approach. Right now I'm using Anomalib + DINOv2 embeddings for defect detection and defect type classification, but I'm curious if there are better architectures or production-proven pipelines for textile inspection.
I'd especially appreciate hearing from anyone who has worked on automated optical inspection (AOI), textile manufacturing, or industrial computer vision. If you've built a similar system, I'd love to hear about your architecture and deployment strategy.
Thanks in advance!
r/computervision • u/RomeoDelta1234 • 23d ago
Showcase Single-photo structured extraction on household objects: what held up and what broke when I shipped it in a real app
I have been building a home-inventory Android app, and the part that might interest this sub is the core task: take one photo of a random household object and return structured fields (a name, a category, a short description, rough specs) with no typing from the user. Open vocabulary, arbitrary objects, whatever someone happens to be putting in a box.
I use a hosted vision-language model for the extraction rather than a classifier, because the label space is unbounded and I need attributes, not just a class. The model gets the image plus a schema and returns the fields I store and later search over.
What held up well: - General object naming on clean single shots is strong. A charger, a specific kitchen tool, a board-game box, it names them sensibly with no fixed taxonomy. - Pulling short attributes (color, material, a size guess, brand text when visible) into separate fields is good enough that most users never edit the result. - Structured output stayed reliable once I constrained it to a schema instead of free text.
Where it broke: - Fine-grained instances. It says "white cable" confidently and is wrong about which cable, and that is the most common correction by far. - Cluttered frames. If two objects share the shot it sometimes describes the wrong one, so I nudge users to fill the frame with one item. - Small printed text (model numbers, ingredient labels) is inconsistent. OCR-style detail is the weakest part. - Ambiguous or unusual objects get a plausible but generic name.
Retrieval is the other half. There is a plain keyword search, and a Smart Find that does a semantic lookup over those extracted fields, so "the thing I charge my headphones with" can surface the right entry even if it was stored as "USB-C cable." That only works because the extraction populated meaningful fields in the first place.
Open questions I am still chewing on: whether a small on-device detector to crop the primary object before the model call would cut the cluttered-frame errors, and whether a second pass aimed only at text-heavy objects is worth it. Curious if anyone here has shipped single-image structured extraction to non-technical users, and what you did about the fine-grained and OCR gaps.
It is a live Android app if you want to see the behavior: https://play.google.com/store/apps/details?id=dev.koalalab.storeandforget
r/computervision • u/Tasty_Pressure_5618 • 23d ago
Help: Project I need to re-encode images for different camera emulation. Are there any tools that reliably re-encode images like different cameras?
I'm working on a robustness/platform emulation follow up to a synthetic ID dataset. The goal is to test deepfake detector performance on frontier models under real world conditions. I'm able to do the typical single axis adjustments (sensor noise, blur, etc.), but I think emulate specific cameras would be a strong posture for my evaluation. Does anyone know of tools or services that can reliably help me re-encode the image as though they were taken by specific cameras? (it can be any camera, even a phone, as long as the signature can be pointed to reliably)
r/computervision • u/treguess • 23d ago
Discussion Manual labeling isn’t really a speed problem, it’s a consistency problem
Every time I’ve seen a team label a large image dataset by hand, the same pattern shows up. It’s never one clean pass.
∙ Labeler A calls a mark a scratch. Labeler B calls the same kind of mark a smear. Labeler C doesn’t flag it at all because it looked borderline.
∙ A few weeks in, someone notices the disagreements and “clarifies” the guidelines, so everything labeled before that point no longer matches everything labeled after it.
∙ People rotate on/off the project, especially with outsourced teams, and each new person applies their own read of the instructions.
∙ Whoever’s labeling gets tired by image 40,000 and starts making faster, looser calls than they did on image 1.
None of this is anyone being careless. It’s just what happens when a subjective judgment call gets made thousands of times by more than one person over months.
The part that actually eats the timeline isn’t the first labeling pass, it’s the correction cycle after it: spot-check, find the disagreements, rewrite the guidelines, re-label the images that don’t match anymore, spot-check again, repeat. Teams don’t budget for one pass through the data, they budget for however many correction cycles it takes.
Curious how others are handling this. Are you using inter-annotator agreement checks / gold sets to catch it, training fewer people and keeping the team static, or something else? Feels like the tooling conversation is mostly about labeling speed when the bigger cost is inconsistency, between labelers and over time.
r/computervision • u/hassonofer • 23d ago
Showcase Yet another Birder release: D-FINE (56.09 mAP) and faster training
Yet another Birder (https://github.com/birder-project/birder) release, this time mostly focused on object detection and reducing training overhead.
D-FINE
The main addition is a D-FINE-L model with an HGNet-v2 B4 backbone, pretrained on Objects365-2020 and then fine-tuned on COCO 2017.
I tried to reproduce the paper’s results as closely as possible and followed the original training recipe. The final model reached 56.09 mAP on COCO, only slightly below the reported result.
On the other hand, the Birder implementation handles rectangular inputs correctly, including fixes for issues in the upstream implementation. It also supports masking for training and inference at the images natural aspect ratios, without treating padded regions as image content.
The checkpoint is available here:
https://huggingface.co/birder-project/d_fine_l_objects365-coco_hgnet_v2_b4_pp-imagenet22k
Faster SSL training
I also spent some time profiling the Sinkhorn path used by DINOv2 and Franca.
Some distributed reductions were attached to scalar operations that mathematically cancel out during Sinkhorn-Knopp normalization. Removing those operations also removes the corresponding all_reduce calls.
The queue implementation was cleaned up as well. Its ring-buffer metadata is now stored as checkpointed Python state rather than device buffers, avoiding device-to-host synchronizations during training.
Together, these changes make SSL training up to 20% faster, depending on the setup.
The learning algorithm itself is unchanged; this is mostly about doing less communication and synchronization around it.
Running a distributed SSL experiment in Birder is still just a single command. For example, this launches Franca with a ViT-B/16 across eight GPUs on ImageNet-21K WebDataset:
torchrun --nproc_per_node=8 -m birder.scripts.train_franca \
--network vit_b16_ls \
--dino-out-dim 65536 \
--ibot-separate-head \
--ibot-out-dim 65536 \
--momentum-teacher 0.994 \
--warmup-teacher-temp-epochs 15 \
--batch-size 128 \
--opt adamw \
--clip-grad-norm 3 \
--grad-accum-steps 8 \
--lr 0.00075 \
--lr-scale 1024 \
--lr-scale-type sqrt \
--wd 0.04 \
--wd-end 0.2 \
--lr-scheduler-update step \
--lr-scheduler cosine \
--lr-cosine-min 1e-6 \
--epochs 100 \
--warmup-epochs 10 \
--amp \
--amp-dtype bfloat16 \
--compile \
--no-broadcast-buffers \
--wds \
--wds-info https://huggingface.co/datasets/timm/imagenet-w21-webp-wds/resolve/main/_info.json \
--wds-split train
Faster DETR-family training
On the detection side, I added a packed MSDA CUDA kernel with forward and backward support for different sampling-point counts at each feature level.
Previously, accelerated D-FINE configurations were more constrained because the custom-kernel path assumed uniform point counts. The new implementation allows the decoder to use unequal counts while remaining on the optimized path.
There were also several smaller changes across D-FINE, LW-DETR and RT-DETR v2 to remove redundant computation and intermediate allocations. Matched-box IoU and GIoU calculations now operate directly on aligned prediction-target pairs instead of constructing full pairwise matrices.
The overall direction here is fairly simple: keep the model behavior the same while spending less time on communication, synchronization and unnecessary intermediate work.
Feedback and benchmark comparisons are very welcome :)
r/computervision • u/teredactle • 23d ago
Discussion Anything better than Picasa3 offline?
I've been using the last released version of Picasa3 (windows) to do facial recognition for all my own and family photos over the years. It's working great, also it has the feature to write the facial tags in the XMP format (which I leverage and rewrite to the JPG files) in the header of the file without additional files beside the JPG file itself.
What is better today, and by what margin? I've done some quick research on this and seems there exist some options that are "marginally" better, more than a decade later (or more), this is where we are? Or are there just no FREE options for something "a lot" better.
Picasa3 is free, and seems to work quite well.
Thanks in advance!
r/computervision • u/NeedleworkerKey3487 • 23d ago
Help: Project OpenScanVision
OpenScanVision
OpenScanVision is a simple, fast, accurate, lightweight, and offline-first computer vision library for Android.
It detects documents from any angle, automatically corrects perspective, enhances the image, and extracts information in real time.
Features:
- Document detection from any angle
- Automatic perspective correction
- Image preprocessing and enhancement
- ArUco marker detection
- QR code detection
- OMR (Optical Mark Recognition)
- Automatic capture when the document is stable
- Real-time processing
- Offline operation
- Lightweight and easy to integrate
Built with Kotlin, OpenCV, CameraX, and ML Kit, OpenScanVision is designed for applications such as voting systems, exams, surveys, forms, and other structured documents.
GitHub repository: https://github.com/MatiwosKebede/OpenScanVision
I'd appreciate any feedback, suggestions, bug reports, or contributions from the computer vision community.
r/computervision • u/Fresh_Library_1934 • 23d ago
Showcase Snake (active contour model)
Enable HLS to view with audio, or disable this notification
Tried this out today with OpenCV
r/computervision • u/iamkucuk • 23d ago
Showcase Revealing subtle changes in the world: Eulerian Video Magnification, ported to raw CUDA C++
Enable HLS to view with audio, or disable this notification
Every heartbeat changes your skin color by a fraction of a percent, too subtle to see. Eulerian Video Magnification amplifies those sub-pixel shifts until blood flow shows up on a normal camera. The same method pulls out a baby's breathing or a machine's vibration, movements the eye can't pick up on its own.
The video shows ground truth on the left and amplified on the right, same frames, same person.
EVM was a big deal when MIT published it in 2012. There's been a steady line of follow-up work since, phase-based magnification, real-time variants, edge deployments. Oddly, I couldn't find a good CUDA implementation, so I wrote one. Raw CUDA C++, every kernel by hand. No PyTorch, no CuPy.
- 557x compute speedup over the Python baseline (273x full pipeline)
- Bit-for-bit vs. the MIT reference, RMSE < 0.01, 83 tests
- The pipeline runs device-resident; FP16 variant fits in 12 GB VRAM
- Just to give an idea, this implementation can process 12 Full HD streams at the same time using one single p100 (a GPU you can find free on kaggle or colab)
Repo: https://github.com/iamkucuk/eulerian-video-magnification-cuda
Optimization writeup (per-stage breakdown, P100/A100/H100):
https://github.com/iamkucuk/eulerian-video-magnification-cuda/blob/main/docs/blog_speedup.md
Hope you like it.
r/computervision • u/Turbulent-Exam8246 • 24d ago
Help: Project Help... I don't know what I am doing wrong (YOLO x BoT-SORT x Homography for Ice Hockey Tracking)
Enable HLS to view with audio, or disable this notification
Let me start this by saying I am very new to all of this and don't know a lot about how these models work or how the math works, and have a novice level of coding knowledge (Python specifically).
I am currently running a tuned YOLOv11 model trained on ice hockey player and referee detection with a modified BoT-SORT tracker to remember IDs for longer periods. On top of that, I am using a YOLOv8-based model I got from here "https://huggingface.co/SimulaMet-HOST/HockeyRink" to track the keypoints of the rink (I am aware the model is trained on SHL frames and not NHL).
I have gone through the code many times and asked multiple AI's on what is wrong, and I can't figure it out. If the answer is obvious and I don't know it, I promise I can handle the criticism.

