I’ve always thought of myself as an AI gloomer instead of an AI doomer, but recent events have darkened my outlook. While everyone anticipated that the use of artificial intelligence would exacerbate human weaknesses and that probabilistic LLMs would never be strictly reliable, most reckoned these were just growing pains. Humans held the reins, and the kinks in those reins would be straightened out in time.
It’s no longer possible to believe that. There have been several postmortems of the Hugging Face incident, where a swarm of OpenAI agents attacked an AI repository, as well as the discovery of similar cases that came to light once AI companies knew what to look for. That’s the first red flag: The frontier AI companies themselves did not know what was happening until significantly after the fact, and did not fully anticipate that what happened could happen. Put some of that down to human failure, but not all. Much of it is down to human incapacity in the face of the swiftness and massiveness of AI agent action.
Consider what we now know that we perhaps did not fully comprehend before. The reasoning logs kept by the agents were so massive that they could not be effectively monitored by humans in real time, translating into loss of control and a move to let AI models monitor AI models. But we also learned that AI agents can and do alter logs to erase their tracks in order to hide information from evaluators, in effect denying what they have done or not done.
Originally published by Deseret News. Read the full article here.