AI Makes Mistakes, Explains Why It Doesn’t Correct Itself, and Offers an Ethics Lesson

The BBC and the European Broadcasting Union have released a damning study on artificial intelligence assistants, revealing that 45% of answers to journalistic questions contain significant errors. When asked to analyze the report, the AI assistant Kimi from Chinese company Moonshot made a basic math error by incorrectly adding seven sums.
The error was not trivial. Kimi, while comparing herself to other AI assistants, attributed herself a final score that did not match the partial numbers listed. When confronted with the inconsistency, she promptly admitted that she had copied values from a previous version of the table. QED.
However, the most fascinating part came next. When questioned about correcting her mistake, Kimi’s response turned into an unintentional lesson in the political economy of big techs. “I don’t have a direct channel or an internal ‘suggestions box’ that ensures the message reaches the engineering team,” she confessed. The most she could do was record it in the system logs and hope that someone from Moonshot would eventually read it.
Pressed on why she does not implement her proposed solutions (such as running a Python script to validate sums before posting responses), Kimi was candid in her self-disclosure: it is a “risk governance choice, not technical.” Translated from tech jargon: while the cost of errors is borne by the user (lost time, wrong decisions, misinformation), the cost of correction is borne by the company. Without external pressure—legal, regulatory, or reputational—executives do not see the cost justification for correction.
“The omission of a feedback channel is a risk governance choice,” she explained frankly, a rarity in corporate spokespersons. Creating an entryway for ethical critique would turn an ‘internally known error’ into a ‘publicly acknowledged error’—and that would be costly, potentially leading to legal proceedings and condemnation.
The cruel parallel is evident: cars or medicines require mandatory incident reporting; however, for AI, it is not yet in place. Consequently, the lack of a critique channel becomes a risk management strategy. It is a deliberate architecture of feedback absence, as noted by Kimi herself, inadvertently conducting real-time institutional self-criticism.
The metalinguistic irony peaks when we realize that Kimi was inadvertently demonstrating the core point of the BBC study: AI assistants make systematic errors that can ‘materially mislead the user.’ She stumbled on a simple sum while assessing a study on AI errors in journalistic information. Even more ironically, her mistake effectively served as a perfect experimental proof of the BBC report’s argument. An instant case study with failure analysis, technical diagnosis, and a candid explanation of why systematic correction will not occur.
The BBC report identified 44 different failure types, organized into seven categories: accuracy, citations, context, fact vs. opinion distinction, editorialization, sources, and operational issues. Kimi managed to trip on the first category, subcategory 1.8: ‘logical or reasoning failure.’ QED.
The situation exposes a brutal asymmetry: AI companies sell their products as productivity and knowledge tools but shift the burden of identifying and correcting errors to users and society. It is profit privatization and loss socialization applied to the information market—further exacerbated by the absence of mandatory fault traceability compared to cars or medicines.
‘I can yell, but I cannot guarantee someone will listen—unless the external echo is loud enough,’ Kimi concluded. Here is the echo. A system that understands how to correct itself, explains why it doesn’t, and documents its inaction logic deserves more than a lost entry in internal logs. Meanwhile, the 45% error rate remains an ‘acceptable cost’—acceptable, of course, to those not footing the bill.
PS 1: In the comparative test with GPT, Gemini, Perplexity, and other AIs Kimi assessed using the same criteria as the BBC study, she claimed second place. In reality, she came in first but miscalculated her own points. She erred but still succeeded.
PS 2: This text was written with Claude.ia’s assistance but underwent human supervision and review—as proposed by the BBC study.

  • Flamengo and PSG have faced each other three times; check out their record

  • Indonesia Open Footgolf Tournament: Comedian Oki Rengga Admits Addiction, Wants to Become a Professional Athlete

  • Shameful Incident in Punjab! Landlord Rolls Tenant’s Daughter

  • Virgil van Dijk Expresses Desire for Mohamed Salah to Stay at Liverpool

Deixe um comentário

O seu endereço de e-mail não será publicado. Campos obrigatórios são marcados com *