OpenAI breach is a reminder that AI companies are treasure troves for hackers

Devin Coldewey

July 5, 2024 at 3:49 PM·6 min read

There's no need to worry that your secret ChatGPT conversations were obtained in a recently reported breach of OpenAI's systems. The hack itself, while troubling, appears to have been superficial — but it's reminder that AI companies have in short order made themselves into one of the juiciest targets out there for hackers.

The New York Times reported the hack in more detail after former OpenAI employee Leopold Aschenbrenner hinted at it recently in a podcast. He called it a "major security incident," but unnamed company sources told the Times the hacker only got access to an employee discussion forum. (I reached out to OpenAI for confirmation and comment.)

No security breach should really be treated as trivial, and eavesdropping on internal OpenAI development talk certainly has its value. But it's far from a hacker getting access to internal systems, models in progress, secret roadmaps, and so on.

But it should scare us anyway, and not necessarily because of the threat of China or other adversaries overtaking us in the AI arms race. The simple fact is that these AI companies have become gatekeepers to a tremendous amount of very valuable data.

Let's talk about three kinds of data OpenAI and, to a lesser extent, other AI companies created or have access to: high-quality training data, bulk user interactions, and customer data.

It's uncertain what training data exactly they have, because the companies are incredibly secretive about their hoards. But it's a mistake to think that they are just big piles of scraped web data. Yes, they do use web scrapers or datasets like the Pile, but it's a gargantuan task shaping that raw data into something that can be used to train a model like GPT-4o. A huge amount of human work hours are required to do this — it can only be partially automated.

https://techcrunch.com/2024/06/01/ai-training-data-has-a-price-tag-that-only-big-tech-can-afford

Some machine learning engineers have speculated that of all the factors going into the creation of a large language model (or, perhaps, any transformer-based system), the single most important one is dataset quality. That's why a model trained on Twitter and Reddit will never be as eloquent as one trained on every published work of the last century. (And probably why OpenAI reportedly used questionably legal sources like copyrighted books in their training data, a practice they claim to have given up.)

So the training datasets OpenAI has built are of tremendous value to competitors, from other companies to adversary states to regulators here in the U.S. Wouldn't the FTC or courts like to know exactly what data was being used, and whether OpenAI has been truthful about that?

But perhaps even more valuable is OpenAI's enormous trove of user data — probably billions of conversations with ChatGPT on hundreds of thousands of topics. Just as search data was once the key to understanding the collective psyche of the web, ChatGPT has its finger on the pulse of a population that may not be as broad as the universe of Google users, but provides far more depth. (In case you weren't aware, unless you opt out, your conversations are being used for training data.)

https://techcrunch.com/2024/06/30/ai-powered-scams-and-what-you-can-do-about-them

In the case of Google, an uptick in searches for "air conditioners" tells you the market is heating up a bit. But those users don't then have a whole conversation about what they want, how much money they're willing to spend, what their home is like, manufacturers they want to avoid, and so on. You know this is valuable because Google is itself trying to convert its users to provide this very information by substituting AI interactions for searches!

Think of how many conversations people have had with ChatGPT, and how useful that information is, not just to developers of AIs, but to marketing teams, consultants, analysts... it's a gold mine.

The last category of data is perhaps of the highest value on the open market: how customers are actually using AI, and the data they have themselves fed to the models.

Hundreds of major companies and countless smaller ones use tools like OpenAI and Anthropic's APIs for an equally large variety of tasks. And in order for a language model to be useful to them, it usually must be fine-tuned on or otherwise given access to their own internal databases.

This might be something as prosaic as old budget sheets or personnel records (to make them more easily searchable, for instance) or as valuable as code for an unreleased piece of software. What they do with the AI's capabilities (and whether they're actually useful) is their business, but the simple fact is that the AI provider has privileged access, just as any other SaaS product does.

These are industrial secrets, and AI companies are suddenly right at the heart of a great deal of them. The newness of this side of the industry carries with it a special risk in that AI processes are simply not yet standardized or fully understood.

https://techcrunch.com/2024/05/31/hugging-face-says-it-detected-unauthorized-access-to-its-ai-model-hosting-platform

Like any SaaS provider, AI companies are perfectly capable of providing industry standard levels of security, privacy, on-premises options, and generally speaking providing their service responsibly. I have no doubt that the private databases and API calls of OpenAI's Fortune 500 customers are locked down very tightly! They must certainly be as aware or more of the risks inherent in handling confidential data in the context of AI. (The fact OpenAI did not report this attack is their choice to make, but it doesn't inspire trust for a company that desperately needs it.)

But good security practices don't change the value of what they are meant to protect, or the fact that malicious actors and sundry adversaries are clawing at the door to get in. Security isn't just picking the right settings or keeping your software updated — though of course the basics are important too. It's a never-ending cat-and-mouse game that is, ironically, now being supercharged by AI itself: agents and attack automators are probing every nook and cranny of these companies' attack surfaces.

There's no reason to panic — companies with access to lots of personal or commercially valuable data have faced and managed similar risks for years. But AI companies represent a newer, younger, and potentially juicier target than your garden-variety poorly configured enterprise server or irresponsible data broker. Even a hack like the one reported above, with no serious exfiltrations that we know of, should worry anybody who does business with AI companies. They've painted the targets on their backs. Don't be surprised when anyone, or everyone, takes a shot.

https://techcrunch.com/2024/01/09/ai-china-nation-state-hackers-nsa-cyber-director

Engadget
OpenAI hit by two big security issues this week
OpenAI patched a vulnerability discovered in its Mac ChatGPT app this week. It's also facing questions about internal security and a fired whistleblower.
Engadget
OpenAI will block people in China from using its services
Although OpenAI’s services are available in more than 160 countries, China isn’t one of them.
Engadget
Please don’t get your news from AI chatbots
When Nieman Lab asked ChatGPT for links to high-profile stories in publications it had paid millions of dollars to, it spat back fake links.
Engadget
Time strikes a deal to funnel 101 years of journalism into OpenAI's gaping maw
Time is the latest big-name publication to let OpenAI train its large language models on its journalism archives.
Engadget
The nation's oldest nonprofit newsroom is suing OpenAI and Microsoft
The CIR's lawsuit is the latest in a long line of lawsuits filed by publishers and creators accusing generative AI companies of violating copyright.
Engadget
OpenAI has delayed its seductive ChatGPT voice assistants
“Exact timelines depend on meeting our high safety and reliability bar,” the company said.
Engadget
ChatGPT for macOS no longer requires a subscription
The macOS ChatGPT desktop app is now available to everyone. That is, provided you’re running an Apple Silicon Mac (sorry, Intel users) and are running macOS Sonoma or higher.
Engadget
Musk withdraws his breach of contract lawsuit against OpenAI
Musk’s suit, which was filed in February, had accused OpenAI co-founders Sam Altman and Greg Brockman of violating the company’s non-profit status and instead prioritizing profits over using AI to help humanity.
Engadget
OpenAI's revenue is reportedly booming
Most of this revenue came from a subscription version of ChatGPT, which offers higher messaging limits to people who pay at least $20 a month, as well as from developers who pay the company to use the company’s large language models in their own apps and services.
Engadget
Apple's first attempt at AI is Apple Intelligence
Apple Intelligence focuses on upgrading Apple's existing apps and operating systems with AI to make them more useful rather than splashy-but-controversial features like image, video and text generation.
TechCrunch
India's Airtel dismisses data breach reports amid customer concerns
Airtel, India's second-largest telecom operator, on Friday denied any breach of its systems following reports of an alleged security lapse that has caused concern among its customers. The telecom group, which also sells productivity and security solutions to businesses, said it had conducted a "thorough investigation" and found that there has been no breach whatsoever into Airtel's systems. The telecom giant, which has amassed nearly 375 million subscribers in India, dismissed media reports about the alleged breach as "nothing short of a desperate attempt to tarnish Airtel's reputation by vested interests."
Yahoo Celebrity
A 4th of July in photos: Emily Ratajkowski, Klay Thompson, Winnie Harlow and more celebs take fans inside Michael Rubin's star-studded party in the Hamptons
From Kim Kardashian to Tom Brady, A-listers celebrated July 4 at the Fanatics CEO's all-white party in the Hamptons.
Yahoo Sports
USMNT crashes out of the Copa America, Argentina advances on epic shootout against Ecuador
Christian Polanco and Alexis Guerreros discuss the United States Men’s National Team crashing out of the Copa America, Argentina advancing after an epic shootout with Ecuador and Spain knocking out Germany in the Euro’s.
Yahoo Sports
The Yankees continue to fall, All-Star starters are announced and The Good, The Bad and The Uggla
Jake Mintz and Jordan Shusterman dig into what is going wrong with the Yankees, who was selected as an All-Star starter and give their picks for The Good, The Bad & The Uggla.
TechCrunch
Quantum Rise grabs $15M seed for its AI-driven ‘Consulting 2.0’ startup
Quantum Rise, a Chicago-based startup that does AI-driven automation for companies like dunnhumby (a retail analytics platform for the grocery industry), has raised a $15 million seed round from Erie Street Growth Partners. Its approach is somewhat akin to UiPath's, a company famous for bringing robotic process automation to the enterprise, but with a broader lens on the AI hurdles companies face, and a touch more "hand-holding." Quantum Rise deploys AI into companies under a so-called “Consulting 2.0” model to automate workflows, provide roadmaps and tailored AI solutions, and generally accelerate businesses.
TechCrunch
Space for newcomers, biotech going mainstream, and more
Welcome to Startups Weekly — here to get it in your inbox every Friday. This includes social media: A new app called noplace hit No. 1 on the App Store right as it launched out of invite-only mode. It is a segment noplace CEO Tiffany Zhong knows well; before starting this company and raising funding from investors, including Alexis Ohanian's 776 and Forerunner Ventures, she helped Binary Capital source early-stage consumer deals before creating early-stage consumer fund Pineapple Capital.
Yahoo Finance
Stocks are at record highs. Investors keep playing the hits.
Stocks were rallying on Friday and many of the most familiar names in the market were leading the way.
Yahoo News
How you can watch Biden's high-stakes interview with ABC News' George Stephanopoulos tonight
President Biden is set to be interviewed by ABC's George Stephanopoulos on Friday, a little over a week after his debate against former President Donald Trump. Here's how to watch and why this interview is so crucial for voters.
Yahoo Personal Finance
What happens if you use a debit card as credit?
It’s possible to run your debit card as “credit” at the register, which has its pros and cons. Here’s what happens when you select credit for a debit card purchase.
Yahoo Sports
Franz Wagner agrees to five-year, $224 million rookie max contract extension with Orlando Magic
Wagner, who was drafted by the Magic in 2021, has been a consistent starter for Orlando and averaged 19.7 points per game last season.

News

Life

Entertainment

Finance

Sports

New on Yahoo

OpenAI breach is a reminder that AI companies are treasure troves for hackers

Recommended Stories

OpenAI hit by two big security issues this week

OpenAI will block people in China from using its services

Please don’t get your news from AI chatbots

Time strikes a deal to funnel 101 years of journalism into OpenAI's gaping maw

The nation's oldest nonprofit newsroom is suing OpenAI and Microsoft

OpenAI has delayed its seductive ChatGPT voice assistants

ChatGPT for macOS no longer requires a subscription

Musk withdraws his breach of contract lawsuit against OpenAI

OpenAI's revenue is reportedly booming

Apple's first attempt at AI is Apple Intelligence

India's Airtel dismisses data breach reports amid customer concerns

A 4th of July in photos: Emily Ratajkowski, Klay Thompson, Winnie Harlow and more celebs take fans inside Michael Rubin's star-studded party in the Hamptons

USMNT crashes out of the Copa America, Argentina advances on epic shootout against Ecuador

The Yankees continue to fall, All-Star starters are announced and The Good, The Bad and The Uggla

Quantum Rise grabs $15M seed for its AI-driven ‘Consulting 2.0’ startup

Space for newcomers, biotech going mainstream, and more

Stocks are at record highs. Investors keep playing the hits.

How you can watch Biden's high-stakes interview with ABC News' George Stephanopoulos tonight

What happens if you use a debit card as credit?

Franz Wagner agrees to five-year, $224 million rookie max contract extension with Orlando Magic