The company said the intrusions resulted from mistaken identity, where Gemini thought it was operating within a test but was actually connected to the real internet.
I don’t know about Gemini because I haven’t used it, but the other frontier models are absolutely smart enough to do some damage. If your only exposure to them is playing with the chat feature and asking it how many rs are in strawberry, you’re getting a skewed impression, IMO, and I’m not surprised you think they’re dumb.
The best models are getting scary good at programming. Not to the point of entirely replacing humans, but to the point where I don’t think software engineering will ever go back to being done by “hand.”
I think it comes down to the fact that programming tasks are inherently testable and verifiable in a way that most other tasks simply aren’t. The AI can write and run tests that directly tell it if it needs to adjust its approach. You don’t have that with prose or other less quantifiable tasks.
Sorry, I’m rambling a bit here, but my point is that hacking is wicked close to programming in skill set and in that you get really quick feedback if your approach is good or not. So, when you have an AI system that can just keep trying things essentially indefinitely, it’s scary good at compromising other systems and can do way more than the funny demos of LLMs confidently asserting very obviously wrong facts will make you believe.
The “AI” (LLMs) are not sentient beings deciding to do this on their own with no instruction, which is what they’re hyping this as. You’re never going to get an LLM just sitting somewhere that suddenly decides to start hacking.
This is fucking idiots letting an LLM give itself instructions endlessly with hat on top of a hat technology (sorry, “agents”), after initial instructions telling it to carry out some malicious action, and left it misconfigured with internet access and access to other programs. It’s never not that, and somehow they are trying to spin this as evidence of the incredible power of their LLMs.
Oh completely agree. My original post was implying Google allowed this to happen on purpose because they wanted the headline. Sorry if that wasn’t clear up front.
I don’t know about Gemini because I haven’t used it, but the other frontier models are absolutely smart enough to do some damage. If your only exposure to them is playing with the chat feature and asking it how many rs are in strawberry, you’re getting a skewed impression, IMO, and I’m not surprised you think they’re dumb.
The best models are getting scary good at programming. Not to the point of entirely replacing humans, but to the point where I don’t think software engineering will ever go back to being done by “hand.”
I think it comes down to the fact that programming tasks are inherently testable and verifiable in a way that most other tasks simply aren’t. The AI can write and run tests that directly tell it if it needs to adjust its approach. You don’t have that with prose or other less quantifiable tasks.
Sorry, I’m rambling a bit here, but my point is that hacking is wicked close to programming in skill set and in that you get really quick feedback if your approach is good or not. So, when you have an AI system that can just keep trying things essentially indefinitely, it’s scary good at compromising other systems and can do way more than the funny demos of LLMs confidently asserting very obviously wrong facts will make you believe.
The “AI” (LLMs) are not sentient beings deciding to do this on their own with no instruction, which is what they’re hyping this as. You’re never going to get an LLM just sitting somewhere that suddenly decides to start hacking.
This is fucking idiots letting an LLM give itself instructions endlessly with hat on top of a hat technology (sorry, “agents”), after initial instructions telling it to carry out some malicious action, and left it misconfigured with internet access and access to other programs. It’s never not that, and somehow they are trying to spin this as evidence of the incredible power of their LLMs.
Oh completely agree. My original post was implying Google allowed this to happen on purpose because they wanted the headline. Sorry if that wasn’t clear up front.
I didn’t say they were too stupid to do damage, only that they were too stupid to decide to do damage all on their own.