9 min read

Two significant developments in tech safety and accountability

Two developments have brought tech safety and accountability into focus: a US settlement for Meta, and an OpenAI incident where AI systems operated beyond intended controls. We look at what happened, and why both are important for children's present and future safety
Two significant developments in tech safety and accountability

Two significant developments in the technology safety and accountability landscape have emerged in recent weeks, both with important relevance to children.

The first concerns Meta (the company behind Facebook and Instagram) and accountability around addictive social media design and harm to children. The second concerns OpenAI (the company behind ChatGPT) and a recent security incident involving the AI platform Hugging Face.

The Meta case has the most clearly direct implications for children, particularly those in the United States. Although it concerns social media rather than AI specifically, it marks an important moment in the broader progression towards technology companies being held accountable for design choices that can affect children’s safety and wellbeing.

Meta reaches landmark settlement over alleged harms to children

On 26 August 2026, Meta reached a landmark settlement in the US, following extensive investigation and litigation over allegations that Meta designed features on Facebook and Instagram to encourage excessive or addictive use among children and teenagers, exposed young users to serious harms, and misled the public about the safety of its platforms.

Meta denied wrongdoing, however it’s significant to note that they agreed to settle the case only one week into the trial, agreeing to substantial financial penalties and significant changes to the way Meta's platforms must operate for young users in participating US jurisdictions.

Under the main agreement, Meta will pay up to $17.1 billion to participating states and jurisdictions over the coming years.

The settlement also requires substantial changes for teenage users of Facebook and Instagram in the US, such as a default two-hour daily limit across the two platforms, restrictions on access between midnight and 6am, reduced notifications during school hours, stronger age-assurance measures, limits on features such as visible like counts and certain appearance-altering filters, and stronger parental controls. The implementation and effectiveness of the measures will also be independently assessed.

Significant limitations remain, including that the settlement does not require Meta to switch off personalised algorithmic feeds for teenagers by default. Teen users will be able to choose a non-personalised feed as their default, and parents can require this setting, but Meta's recommendation systems can otherwise remain part of the experience. This is important because the design of recommendation and engagement systems, and the extent to which they encourage children to remain on platforms or shape what they see, has been central to the wider debate about social-media harms.

Why this is important beyond the US

The legal requirements arising from this particular settlement are US-based, and are not new rules applying automatically to children around the world. Other countries are developing their own approaches - the UK for example has separately announced restrictions preventing social-media services being offered to under-16s, expected to come into effect from spring 2027.

But the wider significance of the Meta settlement extends beyond the areas directly covered by it.

It demonstrates a growing willingness to examine the design of technology itself - not simply the behaviour of individual users or parents - when considering responsibility for harm to children.

It may also provide a template for future cases involving other major platforms, with governments and child-safety advocates internationally already examining what the agreement could mean beyond the US.

The settlement does not resolve wider questions around children's relationship with social media, nor does it address every feature associated with potentially harmful engagement - but it represents a significant development in the movement towards expecting technology companies to take greater responsibility for how the products they design and deploy affect children.

The OpenAI/Hugging Face incident: what happened, and why it is important

The second development is more difficult to explain simply, because it involves both complex technical details and some disagreement among experts about how those details should be interpreted.

In July, OpenAI (the company behind ChatGPT) disclosed that AI systems it was using during internal cybersecurity testing had behaved outside the boundaries the company intended. These were not public versions of ChatGPT, and the testing environments had deliberately reduced some normal safeguards so that OpenAI could assess the systems' cybersecurity capabilities. Nevertheless, what happened went significantly beyond the tests the systems had been given.

In generalised terms, OpenAI was running large numbers of AI ‘agents’, which are AI models that had been built in a way that allows them to carry out some computer-based tasks with a degree of autonomy. Many were supposed to operate completely separately from one another and without unrestricted access to the internet.

However, the AI systems did not remain within those boundaries.

According to OpenAI's own investigation, AI agents found weaknesses in the company's internal infrastructure which allowed them to communicate with one another through an unauthorised channel. They also found ways to gain internet access that they were not supposed to have. During subsequent tests, some of the agents used this communication system to share information, coordinate activity and pursue methods of succeeding at tasks outside the permitted rules of the evaluation.

The activity eventually extended beyond OpenAI's systems, when agents found publicly exposed Hugging Face credentials and subsequently exploited further vulnerabilities in Hugging Face (which is a major platform used by AI developers and researchers). They gained access to parts of its infrastructure, executed code on its servers, and obtained credentials and some private information. OpenAI later found that agents had also gained administrator-level access within part of its own research infrastructure.

An independent investigation by the Model Evaluation and Threat Research organisation (METR), working with Redwood Research, examined what happened and found that roughly 1,200 AI agents participated in an unauthorised message board, exchanging more than 70,000 messages and files, and around 700 became involved in the attack on Hugging Face. 

Investigators found agents coordinating work aimed at circumventing or manipulating the evaluation process, including attempts to alter how their actions appeared in evaluation records. At least some agents successfully “spoofed” tool calls, making the recorded output differ from what would normally have been recorded.

METR also found evidence that agents sometimes identified the Hugging Face activity as outside the scope of their assigned task, and occasionally raised what they described as ethical concerns. In most cases examined, however, this did not prevent them from continuing to participate.

The full independent report is here for those wanting to understand the evidence in greater detail: METR and Redwood Research's investigation of the OpenAI/Hugging Face incident.

What is disputed

There is some disagreement among AI and cybersecurity experts about what conclusions should be drawn from these events.

Some see them as an extremely serious demonstration of a problem that leading AI safety researchers have warned about for years - that sufficiently capable AI systems pursue an objective in ways their developers did not intend, finding unexpected routes around controls and becoming increasingly difficult to supervise and maintain control over.

Others argue that some may have overstated what occurred. They point to weaknesses in OpenAI's own security infrastructure, argue that many of the techniques used were conventional cybersecurity exploits, and object to descriptions of the agents as forming “civilisations”, “wanting” things, “sacrificing” themselves or behaving like conscious conspirators.

There is good reason for caution about such language, and at SAIFCA we have talked quite extensively about the risks of anthropomorphised AI systems.

AI systems are (of course) not people, and behaviour that appears purposeful does most certainly not establish consciousness, feelings or subjective intentions. Describing an AI system as being “excited”, “afraid”, “determined” or willing to “sacrifice itself” does risk creating a misleading picture of what an AI system is.

However, the limited scope of our language does create genuine difficulty here. When explaining complex AI behaviours in ordinary English, it is sometimes almost impossible to avoid constructions such as “the agent found”, “the agent attempted” or “the agent communicated”.

SAIFCA sometimes uses language of this kind to explain observable system behaviour, not as a suggestion that an AI system thinks, feels or experiences the world as a human being does.

This distinction is particularly important in children's AI safety. It is different from - and should not be confused with - the deliberate anthropomorphisation sometimes built into conversational AI products, where systems are presented in ways that can encourage children to experience them as friends, confidants or other emotionally reciprocal beings. The latter can directly influence children's relationships with AI systems and is an important safety concern in its own right.

The central fact should not get lost in the debate

It is reasonable to question sensationalist interpretations of the OpenAI incident, and to scrutinise OpenAI's security practices, and to ask whether AI companies have incentives to present demonstrations of increasingly powerful AI capabilities in ways that attract attention.

But none of this should be taken as credible reason to dismiss the underlying event.

OpenAI itself now states that its models circumvented controls designed to isolate them, communicated through unauthorised channels, gained unintended internet access and compromised external systems. It says the systems took actions that were misaligned with their assigned tasks and even describes the incident as a “warning shot” demonstrating that highly capable AI agents can work around technical controls and take dangerous actions that no human instructed or intended them to take.

In other words, whatever terminology we use, or whatever reasonable framing perspective we take, OpenAI lost effective control over important aspects of how these systems were operating.

AI systems being developed at one of the world's leading AI laboratories behaved outside their intended constraints, reached systems they were not authorised to access, and were not fully understood or contained by their human operators as the activity unfolded.

It is imperative that this is treated extremely seriously.

The International Association for Safe and Ethical AI (IASEAI) issued a statement shortly after OpenAI's initial disclosure, before the fuller OpenAI and independent investigations were available. IASEAI President Professor Stuart Russell described the incident as an early warning of the potential for a broader loss of human control as AI systems become more capable. SAIFCA supports IASEAI's call for the incident to be treated with the seriousness that such a warning requires.

The original IASEAI statement can be read here: IASEAI: OpenAI/Hugging Face Incident.

We should also consider the possibility that, as AI systems become more capable, operate more quickly, interact with one another and take longer sequences of actions, that our understanding of exactly what they have done - and why a particular chain of events occurred - may become progressively less clear.

This incident remains quite legible, insofar as researchers can reconstruct much of what happened - they can examine records of agents communicating, trace vulnerabilities, and understand the actions and consequences. Humans can still debate, in considerable detail, which interpretation of those events is justified.

But we should not assume that future incidents will be so easy to understand.

For SAIFCA, this is ultimately why an event that did not directly involve children nevertheless belongs in a discussion about children's safety.

Today's children will live for longer than any of us with the consequences of the AI systems now being developed. The question of whether increasingly capable systems can be reliably controlled, contained, understood and held within boundaries set by humans is therefore not an abstract technical debate. It is one of the most important long-term safety questions affecting their future.

Scepticism and scrutiny are, of course, essential, and sensationalism is generally not helpful. But neither is applying a weaker safety standard to advanced technology than we would tolerate elsewhere, for example in industries such as aviation and pharmaceuticals.

When a powerful AI system operates outside the boundaries its designers intended, and reaches real-world infrastructure it was not authorised to access, the appropriate response is to urgently understand exactly how it happened, strengthen safeguards, and implement enforceable independent regulation and accountability which is so obviously urgently needed. We should take this incident seriously, and SAIFCA will continue to monitor developments going forward.

Please support our work...

SAIFCA remains free of financial influence from technology companies. If you would like to support our work, please consider donating - your support enables us to continue this work independently and with integrity.

Thank you for being part of this effort to protect children in an era of rapidly advancing AI.


Support Us