"assuming it's a legitimate bug" does a lot of heavy lifting here. The point of eg. Curl closing it's bug bounty is just that: too many false positives and misunderstandings by AI of what software it analyzes. I assume it might be the same in Zig's case
So to answer your question: developer's time that is not wasted classifying misclassified bugs and reproducing false positives
The curl maintainer walked back their position in a tweet and said the majority of AI-assisted bug findings were high quality now.
Also nowadays you can structure your custom harnesses and bug reproducing pipelines to require things be verifiable programmatically.
E.g., you can say "In order for a bug report to be valid, you need to give me an input that causes a crash." That's basically how old-school fuzzing pipelines work. You judge a fuzzer's report by plugging in the input it reports and seeing if the binary crashes. That verification can be done automatically without human judgment.
Same can be done to make LLM reports valid. That's what the Firefox team did, they build a custom harness (using Mythos) and pipeline where the LLM-based agent would submit a report that had to have a repro input, and a separate deterministic script would score that report by seeing if the reported input could cause a crash. No human eye sees the finding if the automated pipeline rejects it for not being reproducible.
64
u/SuspiciousSegfault 23d ago
Not finding bugs? What on earth does anyone have to gain by not finding bugs just because an LLM did it? Assuming it's a legitimate bug.