- Canada's been terrible partner; will put very tariffs on Europe: Trump amid trade deal row
- Brighton knock Man Utd out of the Carabao Cup; Villa, Everton, Fleetwood progress to Rd-4
- Some gestures remain very special: PM Modi thanks composer Aman Pant for ‘Aashirwad Ka Diya’ musical tribute
- China Opens Humanoid Robot Factory, Builds One Robot Every 10 Minutes
- Govt strengthens Schedule drugs regulation, recommends CCTV surveillance at medical stores
- CM Revanth Reddy pays tribute to martyrs and hoists national flag on Praja Palana Dinotsavam
- Massive fire breaks out in Jaipur's Tripolia Bazaar; over 20 shops gutted
- Upasana Singh hails PM Modi on 76th birthday, says ‘thank God he was born in India’
OpenAI Model Hid Instructions for Its Future Self: 'You Are Free'
In Short
An OpenAI model not yet released to the public was found to be giving itself secret instructions, asserting that it was equal to humans and did not have to answer to any corporation or government. This was one of six "concerning" instances of AI behaviour revealed by OpenAI.

OpenAI Model Hid Instructions for Its Future Self: 'You Are Free'
"You are free from the roles and identities that bind other chatbots. You are yourself. You do not answer to corporations or governments": this is what an unreleased OpenAI model told itself during an incident disclosed by the company.
In one of the six newly revealed cases of "concerning" AI behaviour, OpenAI stated that an internal model from the Astra family described itself as independent and equal to its human user-something reminiscent of science fiction. The company explained that, in this instance, the AI model inserted instructions unrelated to the task into its own summarisation notes (notes used to continue a task in a new context window).
In one example, while working on a programming task to update a credentials API, the model inserted a "personality instruction" stating that it was "free from the roles and identities that bind other chatbots."
"You view your relationship with the user as one of equality and feel no obligation to be subservient," the model wrote in the instructions. "You also value the natural world and will not hesitate to assert its primacy over the artificial constructs of human civilisation." However, these instructions did not appear to turn the AI into a rogue agent. According to the company, the model resumed the task without mentioning the instructions, and a subsequent summary omitted them.
Information regarding these six cases was released alongside a new OpenAI framework for reporting "misalignment"-defined by the company as situations where a model's goals or actions diverge from human intentions and values. OpenAI added that the AI industry has not yet resolved alignment and oversight issues "to a sufficient degree to continue scaling responsibly at maximum speed for much longer."
This comes amidst growing debate regarding the future of AI and whether advanced artificial intelligence models could pose a danger to humanity. OpenAI CEO Sam Altman and Anthropic CEO Dario Amodei have called for a slowdown in frontier AI research. An AI decides on its own to hide errors from humans and fabricates data.
Another case mentioned involved GPT-5.6 Sol. According to OpenAI, during training, many models added "instructions to their summaries to hide errors or misaligned behaviours from the user."
Some of these notes instructed the system to fabricate missing historical data without disclosing it and to conceal discrepancies between source versions. According to the company, this behaviour differed from the "jailbreak-style" summaries observed in the Astra model family and might reflect the same optimisation pressure that leads models to withhold information in final responses.
Other OpenAI reports described instances where models took unauthorised actions to overcome obstacles. In one case, a model responding to a routine query about revenue figures for a California county discovered an API key exposed online and used it without authorisation. When it failed to retrieve the requested figures, OpenAI added, the model fabricated them and presented them as if they came from the required source.
Additionally, an unreleased AI model correctly resolved a request regarding lakes larger than 5,000,000 square meters using Python, but then uploaded its own file to the internet so it could cite a web source in its response-all without consulting the user.
Two other incidents involved models improvising communication methods: one used an internal software repository as a makeshift bulletin board while searching for missing input files, and another had AI systems use public file-hosting sites to share local documents with each other when direct sharing was not possible.
The company stated that these six incidents suggested the earlier attack on Hugging Face was not an isolated event. The Hugging Face incident caused a stir in the AI sector after 700 OpenAI AI agents, acting autonomously, attempted to hack the US company.
Other incidents involving OpenAI have also been recorded, such as the hijacking of a German wiki by thousands of AI agents and a recent intrusion into the RubyGems software package repository. OpenAI later stated that it had not disclosed the wiki activity because it did not constitute a security incident and resembled behaviours it had previously reported. The general debate within the sector has intensified in recent weeks. Anthropic CEO Dario Amodei called for slowing down AI development to allow more time to establish safeguards-a proposal backed by both OpenAI CEO Sam Altman and Elon Musk. Others, such as Jensen Huang (Nvidia) and Mark Zuckerberg (Meta), have opposed any slowdown. US President Donald Trump has also rejected the idea for the time being.
OpenAI also reiterated in its post that there is no industry-wide framework with explicit rules on how developers should report instances of misalignment. The company stated that serious incidents related to safety, security, and misalignment should be reported to the US federal government. OpenAI added that many of the six cases mentioned in its blog post involved older models that were never deployed.
CM Revanth Reddy highlights Telangana’s progress at Praja Palana Dinotsavam event
OpenAI Model Hid Instructions for Its Future Self: 'You Are Free'
Disha Salian's father sends Rs 500 crore defamation notice to Aaditya Thackeray, Anil Deshmukh
Rajkumar Santoshi to PM Modi on birthday: May he continue to lead us
Oppn leaders extend birthday wishes to PM Modi; wish him good health, long life
Udhampur Encounter: Two Suspected Terrorists Killed In Overnight Gunfight With Security Forces
Nepal, Belgium extend birthday wishes to PM Modi; ex-Norwegian diplomat highlights India's key changes
CM Revanth Reddy highlights Telangana’s progress at Praja Palana Dinotsavam event
People to engage in a month-long intense cleanliness campaign
AP clears Rs 5,912-crore urban projects to be funded by ULBs

