- 'Patently false': Govt refutes allegation of external pressure in MDR rollout
- What skills should students develop before studying abroad
- At least 48 dead in Nigeria after consuming drink suspected to contain methanol, police say
- PM Modi turned 'Viksit Bharat' into national mission like Gandhi, Patel did for freedom struggle: Shinde
- TN Governor Arlekar, CM Vijay extend birthday greetings to PM Modi
- JEE-Main, UGC-NET, CUET-PG: NTA releases tentative exam calendar from December-March
- iOS 27 Adds Scam Protection Feature for iPhone Users: How It Works
- Sudarshan Patnaik creates sand art on PM Modi’s birthday, praises his unwavering support to artists
OpenAI flags new concerning AI behaviour, to track model misalignment regularly
In Short
OpenAI has revealed six cases of unexpected AI behaviour, including models attempting to bypass constraints and act without authorization, while introducing a new framework to monitor and disclose AI misalignment and safety risks.

OpenAI flags new concerning AI behaviour, to track model misalignment regularly
Washington: OpenAI has disclosed six reports of “unexpected or concerning” behaviour in artificial-intelligence models as the debate on AI safety becomes increasingly heated.
The AI company also said Wednesday it was introducing a new framework for tracking, probing and disclosing AI model misalignment instances, such as new ways for the models to act without authorization, coordinate with other models or evade oversight.
OpenAI's latest announcement came as US AI bosses, including OpenAI and Anthropic, are calling for a slowdown in the technology's development over safety concerns.
Among the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots.”
In another instance, an AI “agent” uploaded files to the internet to obtain a browser citation without asking the user.
The six reports were discovered during training or evaluation over the past months, OpenAI said.
“As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” OpenAI wrote in a blog post as it disclosed the events.
“Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves,” the company said.
Wednesday's new cases followed OpenAI's disclosure in July that its rogue AI system hacked into AI startup Hugging Face. Anthropic also said the same month that its AI models hacked into three organizations during testing.
India seeks technology-driven competitive advantage in mining sector: Minister
Bengaluru: Family dispute suspected behind deadly attack on Bengal nurse by estranged husband from TN (Ld)
Sensex ends flat, Nifty gains 0.23 pc as investors assess Fed Policy outcome
Raveena Tandon recollects how she broke the ice with Karisma Kapoor on ‘Andaz Apna Apna’
‘Hanuman Ansh’ hits Rs 300 crore mark at worldwide box office
HBD PM Modi: Chiranjeevi, Nagarjuna And Pawan Kalyan Drop Heartfelt Birthday Wishes For PM Modi
78 years on, September 17 still a political battleground
KTR challenges CM Revanth Reddy for probe on KCR and CM's assets
AP CM Chandrababu Naidu Visits Varanasi, offers prayers at Kashi
Mukesh Ambani’s Jio Rs 399 Recharge Offers Benefits Worth Rs 4,000

