Picture out of the Google Moffett Place campus when I interned there in 2022.
Published on

Who is to blame?


LLM disclosure: All the words and ideas were tought of and written by me. I used claude code to help edit and clean up in the last rounds of editing. I am sympathetic to Oxide’s RFD on using LLMs as editors.

I’d like to acknowledge Adrian and Sai for going over drafts of this work and follow up discussions.

Yesterday, I watched the Democracy Now Daily Show which covered a particularly concerning story. It was in this CNN report where U.S. military was provided with false intelligence on a Chinese ship by some chatbot. Thankfully, according to the report, additional scrutiny was applied just before the planned operation to intercept the vessel which revealed the error, an operation planned because of the false intelligence provided, an operation that could have escalated into a conflict between the U.S. and China. The first question that came to my mind is this: who is to blame, the individual that filed the report with the help of the chatbot or the chatbot that produced the faulty intelligence at fault? This situation highlights the limitations of the human-in-the-loop system, and, more importantly, the lack of clarity on how we could assign blame.

I shall use chatbots and large language models (LLMs) interchangeably since in most contexts here, they refer to the same system. It is also fair to assume that the chatbot referenced in the report is backed by an LLM.

While it may be obvious to some readers that the responses from LLMs may sometimes be incorrect even when conveyed as some truth, often referred to as hallucinations, I have found that many do not seem to grasp this, or do not care at all. We have seemingly come to associate authoritative language with knowledgeable experts, and models have now employed it whenever we ask them questions. Not to mention the anthropomorphic language tendencies used by such systems which has led to our ‘deskilling’ of human empathy. Regardless to why, the required validation of the responses from LLMs is now often skipped. You see this in the code that people try to merge into open source projects, and, closer to my own experience as an academic, also in the assignment submissions by students for classes and even in paper submissions to academic conferences.

It may be easy to blame the humans that were in the loop that failed to verify the response from the LLM, but the heavy offloading of the blame onto humans has seemingly given companies and organizations the permission to use them imprudently. If a chatbot on a government website, say on tax filing or immigration law, offers incorrect or misleading counsel, is it the individual’s fault if they broke the law due to these responses? I watched and read about a Germany App built on an “AI platform” that could fill out my applications 1, so this question is certainly pressing. There is legal precedent in multiple countries where companies are held accountable for the information provided on their sites via a chatbot, such as the Air Canada case from 2024 and a more recent case in a German regional court that may go to federal court.

Generally, when you contract services from someone to fill out legal forms for you, they can be held accountable if something costly went wrong. However, with LLMs, companies (especially those that develop them), seemingly tried to distance themselves from the very systems they developed and host on their servers (or on a cloud provider). An outrageous case is the countless hacks carried out by swarms of “agents” against other corporations 2, since hacking other companies is considered a felony in the U.S. – I have yet to see a headline of such a charge being brought against the responsible companies.

If such continued lack of scrutiny by governments goes on, there will be an increased lack of incentive for the companies that develop and deploy LLMs to safeguard them, as these systems have been treated as entities without any accountability. I don’t think they should be treated as just tools as we know them. Perhaps we could look at how cars are treated differently from tools you could buy at the corner store, where the manufacturer is still culpable for various faults that extend beyond faulty components such as bad brakes and towards more fundamental issues in its engineering, such as problematic door design.

Perhaps an exaggerated analogy can show my point: if an autonomous car company operated cars that crashed into other cars occasionally, with no fatalities, would it be acceptable even if it happened at rates similar to those at which LLMs hallucinate? Should the company be held accountable if these cars decided to act as a “swarm” and determined that the best way to reach their goal is to drive through buildings? What if there is a human behind the wheel that can intervene? Would it then be the human’s fault if the car crashes? However, notice that in the swarm scenario of the recent “breakouts” from models, they explicitly have no human oversight, so no human behind the wheel. Existing understanding of fault for automobiles would not easily apply here, and we should therefore develop something that is more fitting to how these hypothetical autonomous cars operate. The same should occur with LLMs right now.

Irrespective of how you feel about the above analogy, there is a decent chance that the deployed LLMs have had a significantly higher impact than the autonomous cars. If the LLMs are already deployed for intel in the U.S. military for targeting as reported by CNN above and other sources, it is fair to posit that these systems may have played some role in potential war crimes committed by the U.S. military, regardless of whether the responses provided by the systems were hallucinations or not 3 4. It is not often clear who is to blame.

It does not really make sense to assign the blame to something that is a tool for those with power. Tools that are almost as good as humans, and even worse than humans, have been used to replace human labor for varying reasons. The information age, which started from the advent of computers, has introduced algorithms that could act as oracles for human decisions. Introducing the uncertainty of who we could blame for bad decisions. This has allowed insurance companies, credit reporting companies, and social media companies to make decisions in our lives autonomously at a large scale without a clear entity to blame for bad decisions. The history of accountability for the harms caused by the algorithms and the companies behind them have been a mixed record. This does not mean that we should just ignore the current developments, but as a warning of what impunity has afforded them.

I believe we need to rethink how we treat and assign blame when these systems make mistakes and cause harm. However, if labs can train and serve these LLMs at large cost with impunity, while others deploy them with disregard to security, I fear that consequences may never fall on those that deserve it. Who do we blame had the U.S. carried out its operation against the Chinese vessel? To revisit the above analogy: even if we humans are behind the wheel of a fleet of cars, who do we blame if the fleet rams through the buildings to take us to our destination?

Footnotes

  1. In an ideal world, this would be a godsend if it did more than just online paperwork, but unfortunately a lot is physical. German paperwork sometimes feels harder than research.

  2. A simple search engine search would give you various examples from different companies.

  3. U.S. military drone hitting a school in Iran, Bloomber report and further New York Times analysis.

  4. U.S. strikes on boats around the Caribean and Pacific coasts off of Mexico and South America may be crimes against humanity, as reported by the NYT.