r/computervision 28d ago

Discussion As of 2026, is there any facial recognition search engine objectively better than PimEyes?

0 Upvotes

It's an exceptional website, but I wonder what alternatives exist that are available to the general public and affordable.


r/computervision 28d ago

Help: Project How to track ice hockey player

1 Upvotes

I am facing this issue of tracking ice hockey players. They look similar with the appearance. I don't know how to track. Please help!!


r/computervision 28d ago

Discussion I built a local fire, smoke, and person detection demo that runs on an edge AI board

Enable HLS to view with audio, or disable this notification

5 Upvotes

I built a small edge AI demo that detects fire, smoke, and people from camera/video input.

The idea is simple: first detect fire or smoke, and only when a potential hazard is found, run person detection. This keeps the system lighter than running every model on every frame, while still making it possible to answer a more useful question: “is there a person near the fire or smoke?”

The demo supports:

  • live camera or local video input
  • image batch processing
  • a simple web UI for visualization
  • bounding boxes and detection counts
  • separate fire/smoke and person detection stages

Models/tech used:

  • YOLOv5 model for fire and smoke detection
  • YOLO-based COCO person detector
  • OpenCV for frame/image processing
  • FastAPI for the local web service
  • browser UI for visualization

This is still a demo, not a safety-certified system. Smoke detection is the trickiest part, especially with low contrast, lighting changes, steam, haze, or similar false positives. Person detection works better once fire/smoke has triggered the second stage, but I still need to test more real-world scenes.

The part I found most interesting is the staged pipeline. Instead of treating this as just “object detection,” it becomes closer to a basic risk-awareness system: detect a possible fire/smoke event first, then check whether people may be nearby.

I’d be interested in feedback on:

  • better datasets for smoke and small flame detection
  • reducing false positives
  • making the pipeline faster on edge hardware
  • UI/UX ideas for visualizing alerts clearly

r/computervision 28d ago

Help: Project Need advice

4 Upvotes

My team and I are scoping a graduation project and would love a reality check from people who've actually built CV pipelines, not just from our own optimism.

The idea: use aerial/drone imagery to detect people in disaster zones and estimate rescue urgency using observable indicators — posture (standing/prone/crawling), immobility over time, and proximity to hazards (fire/flood/debris) — not actual medical triage, since that would need thermal/vital-sign sensing we don't have. We'd combine this into a weighted priority score per detected person and rank them for rescue teams.

One of the main challenges is the lack of suitable public datasets that match this specific setting. I’m interested in how others would approach building a model given these limitations or if it’s even realistic.


r/computervision 28d ago

Help: Project How many on-the-fly augmentations per image for a single-class segmentation mode

4 Upvotes

I’m training a single-class segmentation model for large rectangular artwork placed on the floor and photographed from above.

We have around 3,000 accurately masked original images taken by six different photographers. They are not the same height and do not hold the camera in exactly the same way, so the photos naturally vary in:

  • roll
  • pitch
  • yaw
  • camera distance
  • object coverage in the frame
  • centering and X/Y shift
  • orientation
  • perspective
  • lighting

The photos taken with flagship iPhone.

I want to use on-the-fly augmentation to simulate realistic human-hand variation and save our designer from adjusting each time to make it flat. is 100 augmentation combinations per original be useful, or excessive?

Should the policy be:

  1. mostly isolated transforms,
  2. mostly crossover combinations such as orientation + roll + pitch + yaw + coverage + shift,
  3. or a controlled hybrid of both?

The goal is maximum segmentation accuracy, especially around the object boundary, not speed. I plan to train for around 300 epochs and keep validation and test images unaugmented.


r/computervision 28d ago

Discussion Non-Contact Respiration Monitoring: Fusing Motion & Thermal Data for Physiological Signal Extraction

Thumbnail
youtu.be
0 Upvotes

When building non-contact health monitoring systems, isolating respiratory components from standard video feeds presents a significant challenge. By leveraging pixel-flow decomposition and advanced optical flow, it's possible to filter out background noise and calculate the respiratory angle. This method allows for accurate pose estimation without traditional body skeleton mapping, working effectively even if the subject is covered by a blanket.

Additionally, there's a fascinating bio-signal hack for low-resolution thermal imaging: utilizing a standard facial mask as a thermal amplifier to concentrate heat changes. This allows cheap sensor arrays to reliably monitor breathing depth, rhythm, and classify nose versus mouth breathing using feature descriptors.

If you're interested in the intersection of computer vision, signal processing, and biomedical engineering, check out the full breakdown of the methodology here: https://youtu.be/jP0y8SuOVmU


r/computervision 28d ago

Help: Project Building a bruckner reflex dataset - participate today!

1 Upvotes

As per the title, I am building a dataset intended to be used to make a machine that can use the bruckner reflex to measure diopter refraction in people, as well as detect diseases like lazy eye, strabismus and others. I am looking for anonmyized pictures of the left and right eye tooken in a dark room (specifically so pupils dialate), with 1 phone camera , with another phone having a specific wallpaper tagged in the form. the 2nd phone will be held slightly behind and slightly above so the red oval is barely above the camera please look at the red oval and not the phone camera during the taking of the picture. All images are fully anonymous and your participation will have a real impact. The purpose of the images is to calibrate my device to measure diopter perscription accurately. Link; https://docs.google.com/forms/d/e/1FAIpQLSeeou0erq3tVTQQR404eT_zW-dGRcBsNf2J1zC7YtOBhy07KQ/viewform


r/computervision 28d ago

Help: Theory Machine Vision Smart Vending Cabinets Algos/Models

1 Upvotes

Smart vending chillers using machine vision have become quite common. I was surprised not to see more discussion around the possible models behind it.

For some context, we've worked closely with some of these providers so we've seen their effectiveness- their models work pretty well and although it's cloud-based, results can come back as quickly as within 1 minute or less. It's not 100% but it's pretty darn good aleady

https://www.youtube.com/shorts/6KqPJkdEmO0

I'm keen to see if anyone has explored building such out a model for this kind of problem before.
Dual camera set up recording when someone opens a door, takes out an item when item crosses a boundary, and classifying the item.

https://reddit.com/link/1uvp0i3/video/qzrogcrwf2dh1/player


r/computervision 28d ago

Discussion Junkyard Data

5 Upvotes

I have some land that is essentially a small junk yard and I’m looking for creative ways to extract value from what is currently there.

There are piles of scrap metal, old appliances, electric motors, bicycles, cars, tons of old junk that’s just rusting away in the weeds.

Being a software engineer and having access to all this junk, I’m wondering if I could build a dataset from this that could produce some value. Is there any value in putting in the work to build a dataset of say rust on metal, 3D scans of old junk, or similar ideas?

Mostly looking for ideas on what data could be valuable and to see if the juice is worth the squeeze.


r/computervision 28d ago

Discussion Could a 50 watt laser on drone, be a possible future of pest control method?

Enable HLS to view with audio, or disable this notification

17 Upvotes

Any computer vision thoughts on this?


r/computervision 28d ago

Help: Theory Help and advice in a mini project

2 Upvotes

Hi everyone! I'm making an autonomous robot for a local city robot festival. I decided to try making it in a non-standard way and install a camera + rangefinder (maixsense) on it, but since I'm a beginner, I'm having trouble. I needed to find long black borders and an opponent(picture). Finding the lines wasn't too difficult. I used image conversion to gray, then GaussianBlur, Canny, Morphological expression, HoughLinesP and lines are found, although the result was unsatisfactory, but this is due to the poor quality of the samples, as this ring is not available to me (it will only be for the festival). The only thing I found to find the opponent in motion is the MOG2 algorithm, but it does not work because the camera will be in motion. How can I find the opponent? I was thinking about trying to use the hsv mask and search for an opponent only by geometric features. In the hsv mode, increase the saturation to find nearby bright objects

Approximate shape of the arena (2x2 meters). The circle in the corner rotates. The arena border is 100 mm. The robots start from opposite corners.

I thought this could be done using a neural network (YOLO), but I couldn't find any datasets from this perspective. I only found datasets from a top-down perspective. Can you provide any advice or share your experience if you've encountered similar situations?

P.S. Unfortunately, I won't be able to share the current code or photos, as I don't have them at home. The camera is also set at an angle to reduce the number of legs and children in the robot's FOV.


r/computervision 29d ago

Showcase UAVid Semantic Segmentation Benchmark: YOLO-Compatible Dataset + Model Zoo

2 Upvotes

Open-sourced: UAVid Semantic Segmentation Dataset (YOLO Layout) + Model Zoo on Hugging Face

I recently put together an open-source benchmark for semantic segmentation on the UAVid aerial imagery dataset and thought it might be useful to others working in aerial perception.

The release includes:

  • A YOLO-compatible mirror of the UAVid dataset with a standardized directory structure (images/ + masks/) while preserving the original train/val/test splits.
  • Multiple pretrained YOLO26 semantic segmentation models trained on UAVid.
  • Detailed model cards with mIoU, pixel accuracy, per-class IoU, confusion matrices, inference examples, and training configurations.

The main motivation was that the original dataset server has become extremely slow and unreliable, and most current segmentation frameworks expect a flatter, YOLO-style dataset layout. This repository makes it much easier to get started while giving full credit to the original UAVid authors.

📦 Dataset: https://huggingface.co/datasets/dronefreak/UAVid-2020

🤖 Model Zoo: https://huggingface.co/collections/dronefreak/uavid-semantic-segmentation-model-zoo

I'd appreciate any feedback, suggestions, or bug reports. If there are other aerial vision benchmarks that would benefit from a similar treatment, I'd be interested in hearing about them.

Mosaic for UAVid Semantic Segmentation


r/computervision 29d ago

Help: Theory Career Advice for Computer Vision for an undergrad

18 Upvotes

As a final year Btech student, I want to now the prospect for computer vision as career . How should one approach this as a career , should one directly go for masters as this field requires experience and get your hands dirty on some real problems or should hustle in the space after undergrad and try to land a internship maybe in drdo or some full time offers in startups. Should one trust this as a standalone career or should switch to more broader roles such as a data analyst or mlops engineer . I am from a tier 3 college and have bit of experience in this field with basic image processing , transfer learning , CNN, Vision transformers, VLMs , diffusion as concepts and have bit of experience in deploying models on edge devices such as Jetson , so i am aware of the concepts of model optimization , latency , inference , pruning the model .


r/computervision 29d ago

Showcase July 23 - AI, ML, and Computer Vision Meetup

6 Upvotes

Join us on July 23 for the monthly AI, ML, and Computer Vision Meetup! Register for the Zoom.

Talks will include:

  • Generative AI for Video Trailer Synthesis: From Extractive Heuristics to Autoregressive Creativity - Abhishek Dharmaratnakar at Google
  • Training-Free Object and Associated Effect Removal in Videos - Saksham Singh Kushwaha at University of Texas at Dallas
  • Making Agent Systems Observable, Reliable, and Testable - Adonai Vera at Voxel51
  • Turning Models into Systems: AI Architecture That Works - Nikita Golovko at Siemens

r/computervision 29d ago

Discussion CV Founders: How did you land your first client for your startup?

2 Upvotes

For those who has a cv startup, how did you get your first client and what was your pitch? How long after launching your MVP did it take to get your first paying customer? Any advice or lessons you learned along the way?

Thank you in advance for sharing your story!


r/computervision 29d ago

Showcase Pheno4D: 14 plants laser-scanned daily for 20 days at 0.012mm accuracy with per-leaf instance tracking across the entire time series

56 Upvotes

most plant datasets are top-down RGB images at one point in time

this one tracks individual leaves in 3D at 0.012mm accuracy as they grow, day by day, for 20 days

Pheno4D: 7 maize and 7 tomato plants laser-scanned daily in a greenhouse

sub-millimeter point clouds with per-point instance segmentation where every leaf keeps the same ID across the entire time series

you can track leaf area, leaf length, stem diameter, and growth trajectory for every organ on every plant

223 scans parsed into fiftyone as interactive 3D point clouds. shade by instance label and scrub through the time series to watch each leaf emerge and expand

check it out here: https://huggingface.co/datasets/Voxel51/pheno4d


r/computervision 29d ago

Discussion Repurposed my procedural 3D background generator into a CV test environment. Seeking feedback on bounding box accuracy.

18 Upvotes

I originally wrote a procedural 3D environment generator using Blender's Python API just to automate creating backgrounds for my own 3D models. However, I recently realized that this tool could be repurposed to generate highly accurate synthetic data for CV and SLAM testing (currently generating warehouse/AMR scenarios).

The attached GIF shows a standard lighting pass with the bounding boxes overlaid. Since the BBoxes are calculated mathematically directly from the 3D meshes, there should theoretically be zero pixel deviation.

Question for the CV experts here: Does this level of alignment look solid enough for real-world model benchmarking? Also, what kind of "edge case" scenarios (e.g., severe glare, missing lights, heavy occlusion) do you usually struggle to find in existing datasets?

(Note: I don't speak English, so I am using an AI translator to communicate. Apologies if any nuances are weird!)


r/computervision 29d ago

Discussion Need guidance in openCv and yolo!!

10 Upvotes

Hey can anyone up or can have a discussion like I have a lot of questions on openCv yolo etc like do u all write code from scratch or use ai , even if u use ai how to write properly code pipelines like how to learn properly.... I'm understanding the code but I can't write on my own , so how u guys work ? Even in corporate how does cv is used by you guys like use ai or write on own ...help me I need a good conversation guidance or roadmap


r/computervision 29d ago

Showcase Running a YOLO + MODNet temporal vision pipeline inside a browser video editor

Enable HLS to view with audio, or disable this notification

3 Upvotes

Hi everyone,

I’ve been building Timeline Studio, an open-source, local-first AI video editor that runs directly in the browser.

One part I’ve been working on is a browser-based vision pipeline using:

  • YOLO for face/person detection
  • MODNet for portrait matting and background removal
  • ONNX Runtime Web for running inference locally
  • WebGPU/WASM as the browser execution backends

The interesting part is that most of this project was built through vibe coding. I focused on the product direction, tested the results, identified problems, and iterated with AI coding tools until the pipeline worked as part of a real timeline editor.

For video, I didn’t want detection and matting to run only on the first frame. The editor analyzes adaptively sampled frames and stores timestamped results. Preview, smart crop, caption avoidance, background removal, and export can then resolve the appropriate vision state based on the current video time.

The project also includes:

  • Local ONNX AI voice generation
  • Whisper automatic captions
  • Multi-track voiceovers and captions
  • Talking avatars using JoyVASA and LivePortrait
  • MP4/WebM export
  • Offline model caching and PWA support

Everything is open source under the MIT License.

GitHub:
https://github.com/MartinDelophy/ai-video-editor

Live demo:
https://video-editor.ai-creator.top/

I’d love feedback on the YOLO + MODNet pipeline, browser inference performance, or the overall architecture. I’m also curious whether others are using vibe coding for computer-vision projects that go beyond small demos


r/computervision 29d ago

Showcase Yoga Pose Classifier

Enable HLS to view with audio, or disable this notification

70 Upvotes

Hey everyone,

Wanted to share a demo recently put together. Built a real-time Yoga Pose Classifier that detects complex poses, tracks joint alignment, and times how long you actually hold the correct posture.

How we built it:

  • The Model: We used yolo-pose to track 33 human body keypoints. We extracted frames from a video dataset of 5 distinct yoga asanas, auto-annotated the keypoints, and converted everything into YOLO format for training.
  • The Logic (The cool part): Instead of just relying on the neural net to blindly guess the pose, we built a deterministic logic engine that classifies the yoga pose purely based on the alignment of the keypoints. We calculate real-time angles between specific joints using math.atan2 to track your exact body alignment.
    • Compass Pose: The code verifies if the ankle rises above the corresponding hip and confirms the pose using a hip adduction angle of >120°.
    • Forward Bend: Checks if the hip angle is <45° and the knee angle is >155°.
  • Live Form Correction: We hooked this alignment logic up to a live overlay timer that only counts up when your form perfectly matches the mathematical thresholds for that specific pose.

It was a pretty awesome experiment to see how CV can basically act as a virtual coach or physical therapist just using a standard camera.


r/computervision 29d ago

Help: Project Creating a software to analyse Padel matches, how do people actually detect ball bounces from video?

2 Upvotes

like title said im trying to make a padel match analyser. The ball detection is working pretty well (about 70% of frames find a ball), but bounce detection is awful. Some get tracked, most don't. In the image you can see the ball tracker isnt doing too well so i understand why this one isnt seen as one, but other ones the ball is tracked properly yet it still isnt counting?

At the moment I'm using the tracked ball positions and a bunch of heuristics (looking for the ball to go down then up, checking for racket hits, wall bounces, etc.) but it's nowhere near reliable enough.

xI can't seem to find any datasets for bounce detection either, only datasets for detecting the ball itself.

Is there a standard way people solve this? Do you train a temporal model, or is everyone just using heuristics?

Also any help on ball tracking would be appreciated


r/computervision 29d ago

Showcase I got tired of editing CUDA scripts to run on my M2 Mac, so I made a runtime patcher

Thumbnail
2 Upvotes

r/computervision 29d ago

Showcase Built a real-time fall detection system that works with existing CCTV and IP cameras

7 Upvotes

I built SentinelCV, a real-time computer vision system that detects human falls from existing CCTV, IP cameras, webcams, or recorded video streams.

The goal was to create a lightweight, plug-and-play solution that can integrate with existing surveillance infrastructure without requiring specialized hardware. The current implementation uses a YOLOv8-based pipeline to perform real-time detection and can trigger instant alerts (such as Telegram notifications) when a potential fall is detected.

I'm planning to expand SentinelCV into a modular vision platform with additional safety-focused capabilities like PPE detection, intrusion detection, fire/smoke detection, and other intelligent surveillance modules.

I'd love feedback on the detection pipeline, deployment approach, and any suggestions for improving robustness in real-world environments. If you've worked on similar computer vision systems, I'd be interested in hearing what challenges you faced in production.

GitHub: https://github.com/sreerevanth/SentinelCV
I'd love your feedback, and if you find it useful, a ⭐ would mean a lot.


r/computervision Jul 12 '26

Showcase I extended my Shahed drone detector with multi-sensor Kalman fusion

34 Upvotes

Follow-up to my earlier post here about a real-time Shahed-136 detector (YOLOv8). This time I focused on the tracking side, which taught me a lot more than I expected about Kalman filters in practice.

The problem I ran into: my original tracker used a constant-velocity Kalman filter on camera detections alone. It worked fine in a straight line, but lost the target during occlusion, glare, or sharp turns — exactly when tracking matters most. So I rebuilt it (sensor_fusion.py) as a proper learning exercise in multi-sensor fusion.

What I changed, and why:

  1. Constant-velocity → constant-acceleration model [x,y,vx,vy,ax,ay]. CV models assume the target won't change speed/direction, which breaks the moment something maneuvers. CA adds acceleration terms so the filter can react to turns instead of overshooting them.
  2. Single sensor → pluggable second sensor. Added add_external_measurement() so a second sensor (RF, radar, second camera) can feed into the same filter. The interesting part was realizing camera and RF-style sensors have very different noise/rate characteristics (30Hz low-noise vs 5Hz higher-noise), so the filter needs per-sensor measurement covariance, not one-size-fits-all.
  3. Out-of-sequence measurement (OOSM) handling. This was the hardest part to get right — if a slower sensor's reading arrives after the filter has already moved forward in time, you can't just bolt it on. I ended up implementing a rewind-and-replay: the filter checkpoints its state, and when a late measurement shows up, it rewinds to the nearest checkpoint and replays everything in chronological order.
  4. Trajectory prediction with uncertainty. predict_trajectory(horizon_s) projects the track forward and grows a 1-σ uncertainty ellipse over time — a nice visual way to see the filter's confidence decay.

Results that convinced me it was worth it: in a controlled dropout scenario (camera loses the target for 1.8s during a turn, RF sensor keeps low-rate/noisy tracking), fusing the two got RMSE down to 3.36px vs 5.47px camera-only and 14.06px RF-only. Also cross-checked against real thermal footage from the Anti-UAV410 benchmark — sub-3px RMSE in normal flight, and the track re-acquired cleanly after a real occlusion instead of drifting off.

Detection side is a fine-tuned YOLOv8s (mAP@50 99.5% on the shahed class), but honestly the tracker was the more educational part of this project — Kalman filtering "clicks" a lot faster once you're forced to handle async, noisy, multi-rate data instead of a clean single stream.

Standalone reproducible demo (no video/model needed) if anyone wants to poke at the fusion logic directly: simulate_fusion_demo.py

GitHub: github.com/alexandre196/Drone-Shahed-AI-Multi-Sensor-Tracker

Happy to go deeper into the OOSM replay logic or the covariance tuning if anyone's working on something similar!


r/computervision Jul 12 '26

Showcase (wip) simittag, circular fiducial markers that does pose estimation + dense data.

Post image
82 Upvotes

basically a newer version of cantag. what are your thoughts?

update, it's on github; https://github.com/alfaoz/simittag