Rendered at 02:15:15 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
_dwt 9 hours ago [-]
So I think this (AI-written) thing is mostly interesting because of some of the comments, in which commenters lament Cringely putting out this kind of stuff in the lead-up to his death, which they also manage to verify with an obituary.
trollbridge 8 hours ago [-]
Same here. Apparently he really liked AIs, although given his history that's not surprising.
knollimar 7 hours ago [-]
Gives me a kimi k2.5 vibe
trjordan 8 hours ago [-]
What's it audit?
I presume we're still talking about software. Auditing proves what?
Accountants audit that the books are consistent. It's extremely hard to fake consistency and do fraud, especially when there's 3rd parties involved. The edges of the arithmetic get nailed down, because bank payments were sent and invoices paid. So the internal consistency is checkable and valuable.
But software. Are you proving it does what you want? But you didn't write down what you want. Or maybe you're into spec-driven development, or detailed wikis, or whatnot, so you did. Most people aren't, though, and those docs are already out of date anyway.
I think there's something here, because we're all telling our agents what we want the software to be. So maybe you could audit against that. But this isn't the same as accounting.
NoPicklez 47 minutes ago [-]
My thoughts are similarly to SOC 2 type 2, which is that you can do different types of independent audits, whether it be financial, internal controls, cyber security.
But you still need to build the bounds of what you want and what the expectations are.
If you're going to audit software, you need to know what the expectations were when building the software.
Also, AI Assurance is becoming a large field in where these issues are being addressed and it is being performed by independent people.
Part of this auditing story is that it has also introduced lines of defense, so large companies have line 1, line 2 and line 3, all being different degrees of independence.
Basically this article is simply reminding us that you need independent review and verification and you can't trust the product to verify itself. It's nothing new, you need to review the outputs and the mechanics of everything that is important to the organisation. You must trust but verify all outputs and the tools should be able to provide you with that information AND someone must check it.
supern0va 8 hours ago [-]
Where his argument falls apart for me is: why? Audit is intended to address misaligned incentives. It's a form of review done by an independent party, because the incentives are the same for everyone within an organization or with a vested interest in its survival.
The AI doing self-checks through sub-agents (or outsourcing to other models) is a form of review looking for errors, akin to getting a code review for your PRs. You don't need an outside firm or auditor. You need different incentives (a reviewer whose goal isn't to finish the work, but to find problems) and/or different blind spots (trained differently than the model doing the work, which might be more inclined towards that error).
We generally assume that humans can make mistakes and institute mechanisms to flag and correct those mistakes, but that doesn't mean we take that to the level of auditing every discrete action a human at a given job takes.
Likewise for agents, particularly as/if their error rate becomes lower than humans at the given task. They'll get a level of review tied to the importance of their work. And yes, if that's bookkeeping, then they'll probably need a proper audit, just like the humans directing them.
tptacek 10 hours ago [-]
Pangram flags this 100% AI written.
fluidcruft 9 hours ago [-]
It reads like it too. I don't fully understand why a sentence like "It's the signature that costs." makes it into machine output (It costs... what? What does it cost?) But here we are.
axus 9 hours ago [-]
To me it read like a human was pointedly imitating AI , but with a more antagonistic tone. Maybe he asked it to write the article "in his voice".
9 hours ago [-]
trollbridge 8 hours ago [-]
It never was about X. It was Y.
9 hours ago [-]
hapless 9 hours ago [-]
Didn’t Cringely pass away recently?
Hugsbox 9 hours ago [-]
According to a HN post, and basically nothing else that I can seem to find.
Who knows, maybe he's still alive and posted the death notice himself. Dude is/was a compulsive liar. There's still to this day no actual evidence that he genuinely has passed away (unless I'm just unable to find it)
This post is also dated in July, which would have been before he passed away if that post was accurate.
protastus 7 hours ago [-]
The AI verification is not "handled structurally".
The vast majority of my tokens don't go into the authoring -- they go into review iterations. And the problem is that many reviews don't converge, and even the ones that do, can do it slowly, with drift and leave a residue.
Reviews don't converge when the premise is flawed. You can put tripwires to try to catch this, but it's not a robust process. Agents will generally try to act on a flawed premise, especially when they can't foresee the long-term consequences of the axioms under which they are operating.
Reviews drift because the author and reviewer influence each other to depart from the original problem statement. This is also correlated with unbounded scope creep and slop. Again, one can put tripwires to detect this and escalate but agents are extremely creative at finding new ways to ruin your day.
mindslight 8 hours ago [-]
It looks like there is a large error in reasoning - in all of the historic patterns he is giving, the verification is not cheap. Rather what drops in price is the creation, but verification costs real effort, hence the rise of trusted third parties responsible for performing it once and attesting that it was done.
The argument seems like a more appropriate analogy for the slopweb than the dynamics of LLMs themselves.
I presume we're still talking about software. Auditing proves what?
Accountants audit that the books are consistent. It's extremely hard to fake consistency and do fraud, especially when there's 3rd parties involved. The edges of the arithmetic get nailed down, because bank payments were sent and invoices paid. So the internal consistency is checkable and valuable.
But software. Are you proving it does what you want? But you didn't write down what you want. Or maybe you're into spec-driven development, or detailed wikis, or whatnot, so you did. Most people aren't, though, and those docs are already out of date anyway.
I think there's something here, because we're all telling our agents what we want the software to be. So maybe you could audit against that. But this isn't the same as accounting.
But you still need to build the bounds of what you want and what the expectations are.
If you're going to audit software, you need to know what the expectations were when building the software.
Also, AI Assurance is becoming a large field in where these issues are being addressed and it is being performed by independent people.
Part of this auditing story is that it has also introduced lines of defense, so large companies have line 1, line 2 and line 3, all being different degrees of independence.
Basically this article is simply reminding us that you need independent review and verification and you can't trust the product to verify itself. It's nothing new, you need to review the outputs and the mechanics of everything that is important to the organisation. You must trust but verify all outputs and the tools should be able to provide you with that information AND someone must check it.
The AI doing self-checks through sub-agents (or outsourcing to other models) is a form of review looking for errors, akin to getting a code review for your PRs. You don't need an outside firm or auditor. You need different incentives (a reviewer whose goal isn't to finish the work, but to find problems) and/or different blind spots (trained differently than the model doing the work, which might be more inclined towards that error).
We generally assume that humans can make mistakes and institute mechanisms to flag and correct those mistakes, but that doesn't mean we take that to the level of auditing every discrete action a human at a given job takes.
Likewise for agents, particularly as/if their error rate becomes lower than humans at the given task. They'll get a level of review tied to the importance of their work. And yes, if that's bookkeeping, then they'll probably need a proper audit, just like the humans directing them.
Who knows, maybe he's still alive and posted the death notice himself. Dude is/was a compulsive liar. There's still to this day no actual evidence that he genuinely has passed away (unless I'm just unable to find it)
https://www.legacy.com/us/obituaries/name/mark-stephens-obit...
The vast majority of my tokens don't go into the authoring -- they go into review iterations. And the problem is that many reviews don't converge, and even the ones that do, can do it slowly, with drift and leave a residue.
Reviews don't converge when the premise is flawed. You can put tripwires to try to catch this, but it's not a robust process. Agents will generally try to act on a flawed premise, especially when they can't foresee the long-term consequences of the axioms under which they are operating.
Reviews drift because the author and reviewer influence each other to depart from the original problem statement. This is also correlated with unbounded scope creep and slop. Again, one can put tripwires to detect this and escalate but agents are extremely creative at finding new ways to ruin your day.
The argument seems like a more appropriate analogy for the slopweb than the dynamics of LLMs themselves.