- cross-posted to:
- linux@lemmy.ml
- cross-posted to:
- linux@lemmy.ml
Interesting how he mentions LLMs have a 50% error rate, and predicts how no one will pay for a tool with such a poor rate.
Here a comment on this from usernomdeguerre on hacker news (ycombinator.com).
“His” and “GKH” refers to Greg Kroah-Hartmann, who is maintainer of the stable kernel series and the speaker in the video.
I’ve included a few slides into text that i thought were eye-opening to me:
From his Kernel Recipes 2026 slide on Mythos
Mythos's 79 vulnerabilities: 24 - no detail at all "something crashed" 14 - not a bug at all 3 - totally made up data 15 - already fixed in latest release - 11 by others - 4 by anthropic 20 - fixes were needed - 7 "assume a malicious filesystem image" - 2 "assume you can inject a malicious network packet into the middle of the stack" - 2 "NOMMU" - 6 sctp networking issues for untrusted devices - 2 ipv6 minor network issues - 1 gpu driver for local malicious userGHK called this “10 ‘real’ bugfixes”, which to me sounds like there’s a wild hype machine around these companies and uncritical parroting of every press release they make that falls apart when you engage the affected real experts.
It can legitimately find some hard to find bugs but it can quite often come up with stuff for which fix would complicate the code substantially without and practical gains (bugs under very hard to encounter situations etc). So it feels like the biggest problem will arose for setups where people drive this fully autonomously where bug fixes also are implemented autonomously. What would be interesting is whether if this improves when they task a seperate agent to assess practical relevance vs increased code complexity affecting code maintainability and increased interaction between subparts. Both maintainability and practicality however are “long horizon” tasks hard to evaluate without real long term experience.
So it feels like the biggest problem will arose for setups where people drive this fully autonomously where bug fixes also are implemented autonomously.
Completely irrelevant for the kernel and mostly marketing of LLM companies.
Doesn’t mean people aren’t already releasing software pipelined as such or that there aren’t software companies adopting this approach.
Also a completely irrelevant comment would be “I saw a cat today and she was very pretty”, where as you only had to take a single step from that comment to the fact that some of the problems associated to LLMs don’t apply in this context and likely applies to others. So tell me how the fuck is that completely irrelevant. You people sometimes make the most useless comments to appear edgy to a bunch of random strangers in the internet.
Not only that but if by this
mostly marketing of LLM companies
you meant to say “irrelevant to marketing of LLM” companies and not something else, it also does not look like where LLM companies are pushing things towards.
Looks like LLM companies are not only trying to push adoption of their stuff everywhere - wether that might make sense or not - but also are trolling communities which are putting or want to put some sensible limits on that (see discussions around GNOME or KDE).
The fact that a lot of companies are taking crazy pills does not mean that FOSS projects have to do the same.
And the idea that using LLMs for eveything is mandatory is increasingly getting a religigous note. It is absolutely correct to ake a step back - like GKH did here - and access what is helpful or not.
And it is absolutely necessary that communities put limits and boundaries on things that are not helpful.
In case you haven’t realized that is what my first (completely irrelevant) comment exactly advocates. It may be useful with a human in the loop for well-defined tasks. Whether or not it ends up being a boon or a curse on devs remains to be seen. If you have to just sieve through 10 irrelevant bugs to find one very critical one that might be worth it (especially if the LLM can order it by priority reasonably well), could even be a nice training exercise for a more junior coder to get acquainted with the code base. Whether or not these benefits are worth the damage it is causing (nature and society wise) is of course a whole another issue.


