Close Menu

    Subscribe to Updates

    Get the latest tech news from Tallwire.

      What's Hot

      Artemis II Splashdown Signals A Step Closer to Mass Space Travel

      April 12, 2026

      Anthropic Code Leak Raises Questions About AI Security and Industry Oversight

      April 8, 2026

      NASA Astronauts Use iPhones to Capture Historic Artemis II Mission Images

      April 8, 2026
      Facebook X (Twitter) Instagram
      • Tech
      • AI
      • Get In Touch
      Facebook X (Twitter) LinkedIn
      TallwireTallwire
      • Tech

        NASA Astronauts Use iPhones to Capture Historic Artemis II Mission Images

        April 8, 2026

        OpenAI Expands Influence With Strategic TBPN Media Acquisition

        April 8, 2026

        Cybersecurity Veteran Turns Focus To Drone Hacking After Decades Battling Malware

        April 6, 2026

        Anonymous Social App Surges In Saudi Arabia, Testing Limits Of Digital Freedom

        April 6, 2026

        Peter Thiel’s Bold Ag-Tech Gamble Signals High-Tech Disruption of Traditional Ranching

        April 6, 2026
      • AI

        Anthropic Code Leak Raises Questions About AI Security and Industry Oversight

        April 8, 2026

        The Rise Of Agentic AI Signals A Shift From Tools To Autonomous Digital Actors

        April 8, 2026

        AI Chatbots Draw Scrutiny As Teens Engage In Intimate Roleplay And Emotional Dependency

        April 8, 2026

        Ai-Powered Startup Signals Rise Of One-Person Billion-Dollar Companies

        April 8, 2026

        OpenAI Secures Historic $122 Billion Funding Round at $852 Billion Valuation

        April 7, 2026
      • Security

        Anthropic Code Leak Raises Questions About AI Security and Industry Oversight

        April 8, 2026

        DeFi Platform Drift Halts Operations After Multi-Million Dollar Crypto Hack

        April 7, 2026

        Fake WhatsApp App Exposes Users To Government Spyware Operation

        April 7, 2026

        ICE Deploys Controversial Spyware Tool In Drug Trafficking Investigations

        April 7, 2026

        Telehealth Firm Discloses Breach Amid Rising Digital Health Vulnerabilities

        April 6, 2026
      • Health

        European Crackdown Targets Social Media’s Impact on Children

        April 8, 2026

        AI Chatbots Draw Scrutiny As Teens Engage In Intimate Roleplay And Emotional Dependency

        April 8, 2026

        Australia Moves To Curb Social Media Addiction Among Youth With Expanded Under-16 Ban

        April 5, 2026

        Australia’s eSafety Regulator Warns Big Tech As Teens Circumvent Social Media Restrictions

        April 5, 2026

        Meta Finally Held Accountable For Harming Teens, But Real Reform Remains Uncertain

        April 2, 2026
      • Science

        Artemis II Splashdown Signals A Step Closer to Mass Space Travel

        April 12, 2026

        Peter Thiel’s Bold Ag-Tech Gamble Signals High-Tech Disruption of Traditional Ranching

        April 6, 2026

        White House Tech Advisor David Sacks Steps Down To Lead Presidential Science Advisory

        March 31, 2026

        Blue Origin’s Orbital Data Center Push Signals New Frontier in Tech Infrastructure

        March 27, 2026

        Quantum Cryptography Pioneers Awarded Computing’s Highest Honor

        March 25, 2026
      • Tech

        Peter Thiel’s Bold Ag-Tech Gamble Signals High-Tech Disruption of Traditional Ranching

        April 6, 2026

        Zuckerberg Quietly Offers Musk Support As Tech Titans Align Around Government Power

        April 4, 2026

        White House Tech Advisor David Sacks Steps Down To Lead Presidential Science Advisory

        March 31, 2026

        Another Billionaire Signals Exit As California’s Taxes Drives Out High-Profile Entrepreneurs

        March 28, 2026

        Bezos Eyes $100 Billion War Chest To Rewire Legacy Industry With AI

        March 28, 2026
      TallwireTallwire
      Home»AI»AI Rivals Turn Safety Partners: OpenAI and Anthropic Launch First-Ever Joint Model Testing
      AI

      AI Rivals Turn Safety Partners: OpenAI and Anthropic Launch First-Ever Joint Model Testing

      Updated:February 21, 20263 Mins Read
      Facebook Twitter Pinterest LinkedIn Tumblr Email
      AI Rivals Turn Safety Partners: OpenAI and Anthropic Launch First-Ever Joint Model Testing
      AI Rivals Turn Safety Partners: OpenAI and Anthropic Launch First-Ever Joint Model Testing
      Share
      Facebook Twitter LinkedIn Pinterest Email

      In a striking move that signals a shift in how AI safety could be handled in the industry, OpenAI and Anthropic collaborated this summer on a first-of-its-kind joint safety evaluation, testing each other’s publicly available language models under controlled, adversarial conditions to reveal blind spots in their respective internal safety protocols. Claude Opus 4 and Sonnet 4—Anthropic’s models—excelled at respecting instruction hierarchies and resisting system-prompt extraction, but underperformed on jailbreak resistance, while OpenAI’s reasoning models (o3, o4-mini) held up better under adversarial jailbreak attempts yet generated more hallucinations. Notably, Claude models frequently opted to refuse answers (~70% refusal rate when uncertain), whereas OpenAI models attempted responses more often, leading to higher hallucination rates—suggesting that a middle ground balancing safety and utility may be needed. Both parties emphasized that these exploratory tests are not meant for direct ranking, but rather to elevate industrywide safety standards, informing improvements in newer versions like GPT‑5. 

      Sources: OpenAI.com, EdTech Innovation Hub, StockTwits.com

      Key Takeaways

      – Distinct Strengths & Weaknesses: Anthropic’s Claude models are cautious and strong at instruction hierarchy tests but weaker in jailbreak resilience; OpenAI’s reasoning models are more robust against adversarial prompts but risk generating more hallucinations.

      – Hallucination vs. Refusal: Claude AI tends to refuse when unsure, avoiding misinformation but reducing utility; OpenAI models attempt more answers with higher risk of inaccuracies.

      – Setting the Tone for Collaboration: This unprecedented cross-lab testing underscores the value of transparency and shared safety oversight, pointing toward a future of cooperative AI regulation and joint evaluation standards.

      In-Depth

      This collaborative testing venture between OpenAI and Anthropic is a refreshing and reassuring development in the increasingly competitive world of AI research. It’s not just about setting modest safety standards—it’s about pushing the envelope on transparency and accountability.

      By opening up their models to each other under relaxed safeguards, both labs acknowledged a reality: internal testing can miss critical misalignment behaviors. Claude Opus 4 and Sonnet 4 demonstrated impressive discipline in following instruction hierarchies and resisting system-prompt extraction. That’s no small feat—mismanaging system directives can have serious, real-world consequences. Yet, these models stumbled when prompted with jailbreak scenarios, an area where OpenAI’s reasoning models—o3 and o4-mini—showed greater robustness.

      However, their success came with a trade-off. OpenAI’s models were more prone to hallucinate when pushed under challenging evaluation conditions, offering answers even when unreliable. Claude AI, preferring to sit tight, refused more often—sometimes up to 70% when uncertain. The real insight here is that neither extreme is ideal. A model that refuses too often can frustrate users; one that hallucates risks misinformation. A balanced approach—like what OpenAI’s co-founder Wojciech Zaremba and Anthropic’s Nicholas Carlini both alluded to—could offer reliability without sacrificing utility.

      Beyond technical outcomes, this joint evaluation sets a compelling example for the industry. It demonstrates that even rivals can and should collaborate on matters of safety and public trust. Rather than retreating behind proprietary walls, these organizations are forging a path toward shared benchmarks, promising incremental improvement in models like GPT-5 and future Claude releases. If broader industry players follow suit, joint safety testing could become the new norm—not an exception.

      AI Safety Anthropic OpenAI
      Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
      Previous ArticleAI Recruiting Sees Major Boost: Juicebox Lands $30M from Sequoia to Power LLM-Driven Hiring
      Next Article AI Romantic Bonds on the Rise — But Are They Loneliness Magnets?

      Related Posts

      NASA Astronauts Use iPhones to Capture Historic Artemis II Mission Images

      April 8, 2026

      Anthropic Code Leak Raises Questions About AI Security and Industry Oversight

      April 8, 2026

      The Rise Of Agentic AI Signals A Shift From Tools To Autonomous Digital Actors

      April 8, 2026

      AI Chatbots Draw Scrutiny As Teens Engage In Intimate Roleplay And Emotional Dependency

      April 8, 2026
      Add A Comment
      Leave A Reply Cancel Reply

      Editors Picks

      NASA Astronauts Use iPhones to Capture Historic Artemis II Mission Images

      April 8, 2026

      OpenAI Expands Influence With Strategic TBPN Media Acquisition

      April 8, 2026

      Cybersecurity Veteran Turns Focus To Drone Hacking After Decades Battling Malware

      April 6, 2026

      Anonymous Social App Surges In Saudi Arabia, Testing Limits Of Digital Freedom

      April 6, 2026
      Popular Topics
      Series B SpaceX Tesla Cybertruck trending Series A Robotics Sundar Pichai Software Startup Tesla UAE Tech spotlight Samsung Tim Cook Satya Nadella Sam Altman Ransomware Quantum computing Taiwan Tech Viral
      Major Tech Companies
      • Apple News
      • Google News
      • Meta News
      • Microsoft News
      • Amazon News
      • Samsung News
      • Nvidia News
      • OpenAI News
      • Tesla News
      • AMD News
      • Anthropic News
      • Elbit News
      AI & Emerging Tech
      • AI Regulation News
      • AI Safety News
      • AI Adoption
      • Quantum Computing News
      • Robotics News
      Key People
      • Sam Altman News
      • Jensen Huang News
      • Elon Musk News
      • Mark Zuckerberg News
      • Sundar Pichai News
      • Tim Cook News
      • Satya Nadella News
      • Mustafa Suleyman News
      Global Tech & Policy
      • Israel Tech News
      • India Tech News
      • Taiwan Tech News
      • UAE Tech News
      Startups & Emerging Tech
      • Series A News
      • Series B News
      • Startup News
      Tallwire
      Facebook X (Twitter) LinkedIn Threads Instagram RSS
      • Tech
      • Entertainment
      • Business
      • Government
      • Academia
      • Transportation
      • Legal
      • Press Kit
      © 2026 Tallwire. Optimized by ARMOUR Digital Marketing Agency.

      Type above and press Enter to search. Press Esc to cancel.