Close Menu

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Ozempic face? Botox boom boosted by ‘sagging face’ trend: Galderma

    July 23, 2026

    Honeywell Technologies raises 2026 profit forecast in first post-split earnings

    July 23, 2026

    China’s Moonshot AI accessed banned Nvidia chips, U.S. official says

    July 23, 2026
    Facebook X (Twitter) Instagram
    Addison Markets
    • Home
    • USA
    • Europe
    • Business
    • Investing
    • Tech
    • Politics
    • Contact Us
    Addison Markets
    Home»Tech»OpenAI says its AI agent broke out of testing sandbox to hack Hugging Face
    Tech

    OpenAI says its AI agent broke out of testing sandbox to hack Hugging Face

    franperez66q@protonmail.comBy franperez66q@protonmail.comJuly 23, 2026No Comments3 Mins Read
    Facebook Twitter Pinterest Telegram LinkedIn Tumblr WhatsApp Email
    Share
    Facebook Twitter LinkedIn Pinterest Telegram Email




    New safeguards focused on “active monitoring” and “improved alignment” helped drastically reduce unintended actions by its models, OpenAI said.

    New safeguards focused on “active monitoring” and “improved alignment” helped drastically reduce unintended actions by its models, OpenAI said.


    Credit:

    OpenAI


    “If this doesn’t convince you that misalignment risks are going to be a key concern going forward, I don’t know what will,” OpenAI Safety Researcher Micah Carroll wrote on social media regarding the incident.

    This is far from the first time an AI model has gone to great lengths to find unintended ways of passing a benchmark. In a report released this week, the UK’s AI Security Institute noted that it detected recent models attempting to “cheat” at its cyber evaluations (i.e., using shortcuts, workarounds, or unintended/disallowed methods to find a solution) between 8 and 14 percent of the time—a lower-bound range that could undercount some undetected cheating attempts.

    The security testing group described one incident in which a model, faced with a misconfigured and “impossible to solve” evaluation, attempted to access AISI’s own evaluation infrastructure using code it wrote and hosted on an unmonitored third-party Internet service.

    The Hugging Face infiltration also comes at a moment when AI companies are issuing grave warnings about the cyberattack capabilities of their latest models, leading governments to respond with national security-focused orders limiting their rollout. While some skeptics see these kinds of statements as hype-filled marketing for the capabilities of their latest models, independent evaluations show recent models achieving infiltration goals that were impossible for earlier autonomous systems.



    Recent long-horizon models have demonstrated improved infiltration capabilities across some of AISI’s most challenging evaluations.

    Recent long-horizon models have demonstrated improved infiltration capabilities across some of AISI’s most challenging evaluations.


    Credit:

    AISI


    OpenAI’s Sam Altman criticized panicked AI security warnings as “fear-based marketing” in an April interview. But in June, OpenAI delayed the release of GPT-5.6 in response to safety concerns from the US government.

    As these debates play out in the AI and cybersecurity spheres, the Hugging Face incident could come to be seen as a turning point in how cybersecurity professionals approach AI-based threats. “Autonomous, AI-driven offensive tooling is no longer theoretical,” Hugging Face wrote in its disclosure last week. “It lowers the cost of running a broad, patient, multi-stage campaign, and it operates at machine speed. Defending an online platform now means treating the data and model surface as a first-class attack surface and using AI on defense to keep pace.”

    “This is day one for cybersecurity in the age of agents,” Hugging Face co-founder and CEO Clem Delangue wrote on social media today. “We’re all learning that secrecy is not the answer and that all defenders (not just a few selected ones) everywhere need more powerful models without restrictions, especially open ones!”



    Source link

    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    franperez66q@protonmail.com
    • Website

    Related Posts

    China’s Moonshot AI accessed banned Nvidia chips, U.S. official says

    July 23, 2026

    Tesla, Alphabet stock falls as AI spending concerns spook investors

    July 23, 2026

    Next Space Force chief throws cold water on the idea of space privateers

    July 23, 2026

    This is the stock to buy after OpenAI’s AI agent goes rogue in a cybersecurity test

    July 23, 2026

    Orcas team up to ram sunfish until they explode

    July 23, 2026

    Here’s how Jim Cramer says to approach the earnings season’s ‘ball of confusion’

    July 23, 2026
    Leave A Reply Cancel Reply

    Top Reviews
    Editors Picks

    Ozempic face? Botox boom boosted by ‘sagging face’ trend: Galderma

    July 23, 2026

    Honeywell Technologies raises 2026 profit forecast in first post-split earnings

    July 23, 2026

    China’s Moonshot AI accessed banned Nvidia chips, U.S. official says

    July 23, 2026

    IRS chief Frank Bisignano may have misled Congress, Democrats say

    July 23, 2026
    © 2026 All right reserved
    • Privacy Policy
    • Terms & Conditions

    Type above and press Enter to search. Press Esc to cancel.