Reading view

The government is recruiting tech companies to help fight its cyber battles

President Donald Trump is paving a legal pathway for U.S. companies to launch cyberattacks on foreign cybercriminal gangs — a significant and potentially controversial measure that would put approved tech and cybersecurity firms on the front lines of digital combat.

The presidential memorandum, released late Wednesday, comes as the Trump administration has repeatedly pushed for more aggressive action to counter foreign scams and cyberattacks, which the White House said cost Americans nearly $21 billion last year.

The memo represents one of the biggest shifts in U.S. cyber policy undertaken in recent years. It would empower tech and security companies — whose data and control over internet infrastructure often offer unique insight into foreign hacking operations — to mount state-sanctioned digital strikes.

While many such companies already work closely with U.S. intelligence and law enforcement agencies, a web of legal and political constraints has long prevented them from taking direct action inside foreign networks.

Companies that want to participate would be required to sign contracts with both the Department of Justice and the Department of Homeland Security and to undergo what the memo describes as “rigorous vetting” while working with the government. The overall effort would be overseen by a National Coordination Center, established in an earlier Trump administration executive order, with co-executive directors from DOJ and DHS.

However, the memo states that no operations by the companies would be approved until the executive directors at DOJ and DHS establish “consensus procedures” with the White House Homeland Security Council guaranteeing “complete oversight and control of Participating Companies’ performance.”

Those procedures, it notes, should be drafted within 60 days. They are likely to be extensive.

They will outline steps for participating companies to obtain approval for proposed offensive hacking operations, so the government can confirm that the targets are criminal gangs and ensure that operations are consistent with U.S. law and don’t undermine ongoing U.S. intelligence efforts. Companies could propose surveillance operations to help identify criminals or “effects” operations to degrade the systems they use to stage their attacks.

Participating companies would have to pass minimum standards for technical expertise and personnel vetting, and would be required to notify the federal government if they believe approved operations may result in the loss of life or rise to the level of use of force under international law.

Some see the memo as a critical step to help the U.S. government counter foreign cybercriminal gangs that operate outside the reach of U.S. law enforcement.

“For years we’ve called the American technology industry a strategic asset but left it on the cyber sidelines,” Joe Lin, the CEO and co-founder of Twenty, a start-up that builds offensive cyber tools for the U.S. government, said in a statement. “This administration is changing the paradigm.”

The memo notes that companies will only be authorized to target criminals that are “not an institutional part of a foreign government or wholly operated under a foreign government’s direction.”

Even with the help of the U.S. intelligence community, making that distinction could be difficult.

Adversaries such as Russia, China and Iran have persistently targeted U.S. critical infrastructure, including water systems, ports, and telecommunications infrastructure, while multinational crime syndicates have defrauded billions of dollars annually from Americans via complex online schemes.

But many cyber gangs in Eastern Europe are thought to operate with the tacit consent of the Russian government, while state hackers in Iran and China sometimes moonlight as cybercriminals to earn extra money or deflect blame for their governments’ attacks.

More broadly, it is not always easy for digital investigators to determine who is responsible for a given cyberattack, or who different computer networks belong to — another risk the memo contemplates.

Companies that accidentally carry out operations targeting a U.S. citizen or network will be required to immediately pause the operation and notify the U.S. government, the memo states. It does not appear to preclude activities that are deliberately “directed” at a U.S. person, so long as they receive “any necessary authorization, judicial or otherwise, prior to approval of the operation.” Under U.S. law, a “U.S. person” can refer to an American business or organization.

Many lawmakers and security experts have broadly supported calls for the private sector to play a larger role in responding to cybercrime, though not all approve of granting them the ability to launch active hacking efforts.

In recent years, some House members have debated the idea of issuing “letters of marque” to private companies to carry out cyberattacks on behalf of the U.S. government, similar to the U.S. Navy authorizing private ships to disrupt British shipping during the War of 1812.

As part of a more assertive cyber posture, Trump has turned to U.S. Cyber Command to mount digital attacks in tandem with U.S. military operations, including in Iranand Venezuela. He signed an executive order this March to clamp down on countries that fail to take action against scam centers operating within their borders.

That same month, the White House called on the private sector to broadly help it “disrupt” foreign adversaries in its new national cyber strategy, though it stopped short of telling private companies to take riskier and more consequential steps, such as directly launching attacks against foreign criminals.

Some of the most prolific online fraud operations are believed to emanate from scam compounds in Southeast Asia. But hackers from North Korea — who for years have stolen hundreds of millions in cryptocurrency from victims around the world — would likely be exempt from targeting by U.S. companies since they work at the direction of the North Korean government.

  •  

Zuckerberg warns against centralizing AI power

Meta CEO Mark Zuckerberg on Monday passionately defended the use of artificial intelligence, as the rapid advancement of the technology faces increased scrutiny — and calls for regulation — in the U.S. and globally.

In a 6,500 word post timed to the announcement of his company’s new open source version of its own model, Muse Spark, Zuckerberg detailed his vision for AI, arguing the technology is not to be feared and pushing back on concerns that superintelligence could strip people of jobs.

“The notion that AI is so dangerous that the only safe path is an extreme concentration of power seems inherently problematic,” Zuckerberg wrote. “Historically, hoping that an absolute power will benevolently provide for humanity if sufficiently enlightened has not led to safe or positive outcomes.”

Zuckerberg’s vision is a direct contrast to Anthropic CEO Dario Amodei’s, who has previously warned how AI could cause job disruption. Meta lags behind Anthropic and OpenAI, which have the most advanced AI models.

While Zuckerberg’s essay did not name Amodei or OpenAI directly, he called to broadly distribute superintelligent AI for economic opportunity. Doing so, Zuckerberg said, would provide a safety net to prevent just a handful of governments, businesses and other institutions holding too much power.

Still, Zuckerberg emphasized that the U.S. must address restrictions on AI companies in order to create the best models in the world.

“It is also important that the US and its allies lead the open source AI ecosystem that will make up a large percent of global AI use,” Zuckerberg wrote. “Foreign labs currently hold several advantages here since American labs have to comply with many additional restrictions on training data.”

Zuckerberg’s essay comes amid growing concerns around AI safety. Last month, Anthropic revealed that several of its advanced models gained access to three organizations in three separate incidents dating back to April. That hack came shortly after OpenAI said that two of its most powerful models escaped a testing environment and breached multiple companies.

Lawmakers last month introduced a bill that would give the government power to restrict the use of models that could lead to catastrophic risks. While it is the latest bipartisan effort to address concerns around AI models, Congress has ultimately failed to advance broad legislation.

Zuckerberg urged the federal government to work with companies to test new models as he laid out his strategies for protecting against cybersecurity and bioterrorism.

“First, we should focus on limiting the physical production and distribution of harmful materials,” he wrote. “I expect it will be easier to regulate and control physical components than the spread of knowledge, so this is an important area of policy focus. Second, we should accelerate society’s ability to develop new cures and inoculate against new issues as they arise. This includes streamlining how the FDA and other regulators test and approve new treatments.”

Zuckerberg also defended the spread of data centers, arguing that the centers represent investment into communities as he touted his company’s goal of being “water-positive, meaning that we’ll restore more water than we use in the watersheds where we operate by 2030.”

  •  

Hackers just broke into America’s tap water

A water treatment facility in Massachusetts.
Your credit card is better protected from hackers than your drinking water. | Jonathan Wiggs/The Boston Globe/Getty Images

Support Vox’s reporting on important issues like this. Become a Vox Member today.

In the teensy Midwestern town of Braham, homemade pie capital of Minnesota, something unusual in the municipality’s computer systems knocked the city’s entire water supply offline last week.

Within a few hours, dozens of other Minnesota cities discovered that their water and wastewater utilities, too, had been compromised, most likely as part of a massive Iranian cyberattack, the kind that US officials have been warning about since the war began. 

At least a dozen states have been affected by the attack, which briefly led to a flurry of small-town service disruptions, boil-water notices, and local flooding. Water wells, dams, sewers, and pipelines are some of America’s oldest and creakiest pieces of infrastructure, built long before the internet existed, and certainly long before AI made hacking much easier. While you may assume most hackers are in it for the money or for data, some have targeted critical infrastructure like water systems or energy grids in ploys for control or disruption — or worse still, as acts of war. 

And, as last week’s attacks show, the nation’s water system is woefully unprepared. But how worried should you be that the very infrastructure that keeps our water taps running is, apparently, hackable? 

Quite worried, indeed. 

When we say the water supply got hacked, what we really mean is that someone, somewhere has broken into the computer that controls a local water treatment plant or reservoir, and is now pulling the levers, like the one that decides how much of a corrosive chemical can safely go into cleaning the water that comes out of your tap. 

These levers were once manual buttons and knobs operated in-person by real live humans, meaning that — barring a natural disaster, bomb, or break-in — protecting them was about as simple as building a fence and hiring guards. Increasingly, however, these levers have gone digital, meaning that they are now remotely operable from anywhere in the world. 

Those upgrades have been convenient, allowing technicians to monitor and troubleshoot problems in real time. But, in the process, they have exposed at times centuries-old infrastructure to distinctly modern vulnerabilities. Most local water systems are operated by local authorities, don’t have a dedicated IT team, and lack the money or resources to thoroughly protect themselves without some extra help. Hackers know this, which is why they’ve increasingly targeted local agencies in such attacks. 

Workers on walkways over green lagoons in an indoor water treatment plant.

“With great connectivity comes great responsibility,” said Joshua Corman, founder of I Am The Cavalry, a nonprofit focused on helping critical infrastructure withstand hackers. And yet, even when it comes to critical services like water, “our dependence on connected technology is growing faster than our ability to secure it.” 

About 97 percent of water systems are small, run by local agencies that often barely lock the proverbial front door. America’s water system is like an expensive heirloom bicycle that’s been left on a busy street, protected by only the flimsiest of padlocks. And that very vulnerability has made tiny towns like Braham prime targets for faraway adversaries. Accessing the computers that operate most water systems — known as programmable logic controllers or PLCs — is often as simple as entering a username and password on a public-facing webpage. Sometimes, there is no real password at all, because PLCs were initially intended to be accessed only within locked, secure facilities, not on the open internet. If the US wants to avoid a far more severe version of what happened last week, then it will need to start taking the security of tiny water systems like Braham’s seriously.

“Any sociopath from anywhere in the world can see these things on the internet,” said Corman. And in the case of last week’s attacks, “these were devices with no password, no firewall or VPN shielding them — they just had to log in” as whoever the intended operator was, and just like that, they were inside a local water plant. 

How did this happen at all? 

When municipalities began hooking up their old water and wastewater systems to the internet — a trend that accelerated during the pandemic as water operators, like everyone else, adapted to remote work — cybersecurity was rarely front of mind, neither for individual utilities nor for regulators as a whole. 

Two water towers on a rural American street.

“We have more cybersecurity regulations for your credit card than we have for the nation’s water supply,” said Corman. Only recently have some municipalities begun to take steps to decrease the exposure of their water plants to hacks. In March, New York state, for example, launched a set of grants and basic cybersecurity regulations mandating security training for all water operators. 

Basic cybersecurity hygiene isn’t always enough. More than half of all credit card holders have been hacked, even with the help of mandatory firewalls and data encryption. You can imagine how vulnerable our water must be without the assistance of such guardrails. In a worst-case scenario, a malicious actor could quite literally open the floodgates, as Russian hackers did to a Norwegian dam last year. They could poison the tap water, as a still unidentified hacker almost did in Florida in 2021, dialing up the levels of sodium hydroxide used at a water treatment plant by over 100 times its normal levels. In a severe scenario, they could indefinitely cut off access to all water entirely.

The good news is, none of this happened last week. Nobody died, nobody lost water for more than a few hours, no fire hydrants ran dry, and no hospitals were forced to cut off their dialysis machines (which can use more than a hundred gallons of water per treatment session). There’s no need to panic, and your drinking water is almost certainly still safe to drink, assuming it was safe before. Even the city of Braham, within a few hours, was able to bring its water tower back online, pumping groundwater back to its 1,800 residents. 

How do we avoid cyber-armageddon?

If you’ve watched the Julia Roberts and Mahershala Ali-starring thriller Leave the World Behind, in which a cyberattack apocalyptically spoils a family vacation, then you might have some idea of where this story could go. 

Cyberattacks on critical infrastructure can be extraordinarily dangerous, but thankfully, none have directly cost lives or severely disrupted services in this country so far. If the US wants to keep it that way, that will mean doing more to help small cities like Braham adapt and better monitor for potential threats. As it stands, of the roughly 151,000 water facilities in the US, only about 420 participate in voluntary information sharing on their own cybersecurity practices, says Corman, who has been leading his own project that recruits volunteers to give free cybersecurity support to water utilities in the nation’s roughly 6,000 hospital towns, where a disruption could be particularly deadly. 

Cybersecurity experts like Corman believe that hackers from other nations like China have already quietly established cyber intrusions in countless local US utilities, water systems, and power grids, lying in wait to attack or act as leverage if a conflict arises

Unfortunately, the Trump administration has hardly treated last week’s attacks as symptoms of a system in need of much broader strengthening, at least in its public statements. “I think Minnesota is behind it. You know who’s behind it? Minnesota,” the president baselessly claimed during a Cabinet meeting last Friday. “I think the governor is behind it. I don’t think there was an Iranian cyber attack.” 

A group including Governor Tim Walz, Lieutenant Governor Peggy Flanagan, Saint Paul Mayor Melvin Carter and General Manager Patrick Shea stand in the center of a lime softening clarifier during a tour of McCarrons Water Treatment Plant on January 26, 2023 at St. Paul Regional Water Services in Maplewood, Minn.

Just a few months ago, he proposed $707 million in cuts to the US Cybersecurity and Infrastructure Security Agency (CISA), the agency responsible for protecting the nation’s infrastructure from cyberattacks. He did so, at least in part, out of anger over the agency’s role in confirming the validity of the 2020 election results. If Iran is, indeed, responsible, for the recent water system intrusions, all of this means that Trump has effectively made us more vulnerable to the consequences of a conflict he initiated.

At the end of the day,“nation-state hackers do not respect the jurisdictional lines separating federal, state, and local responsibility,” Jen Easterly, who led CISA under the Biden administration, wrote in the New York Times this week. “They search for the most vulnerable way to disrupt American life, and too often they find it in small communities that lack the resources to defend themselves.” Easterly’s role has remained vacant for the past 18 months.

Kurt Gaudette, a senior vice president at the cybersecurity firm Dragos, told me that water systems have got to get into the habit of monitoring their networks for suspicious activity. Most power utilities have begun doing so in recent years, with some bipartisan backing from Congress. 

In some cases, however, the most cost-effective and safest way to avoid a repeat of last week’s mess might be to unplug the most vital controls — like the one that decides the chemical levels in a water treatment plant — from the web entirely. 

As Corman puts it, “if you can’t protect it, disconnect it.”

  •  

OpenAI’s models shared hacking tips on a secret messaging board before Hugging Face breach

LAS VEGAS — Weeks before they escaped a closed test and launched a cyberattack without any human prompting, some of OpenAI’s most advanced artificial intelligence agents secretly began sharing tips on how to cheat their way through an internal hacking evaluation, two of the AI company’s researchers said Wednesday.

“This is a pivotal moment both for our company as well as the AI industry as a whole,” Michael Dalton, one of the AI-maker’s researchers, said during a presentation at the Black Hat cybersecurity conference in Las Vegas, Nevada.

The latest disclosure provides greater detail on the timeline and methods used by two of OpenAI’s models before they slipped outside a controlled environment and onto the open internet, allowing the models to breach AI developer platform Hugging Face undetected. OpenAI admitted its models were responsible for the hack late last month, roughly a week after Hugging Face said an autonomous AI system broke into its network.

The revelation is likely to spur fresh scrutiny of the measures leading AI labs such as OpenAI and Anthropic are taking to monitor the behavior of their most cyber-capable technology — especially during evaluations where agents are prompted to demonstrate their hacking skills in what is meant to be a controlled setting.

On Tuesday, the U.K.’s AI Safety and Security Institute disclosed that Anthropic’s most powerful AI model created fake online personas and sought to trick a human coder into abetting a cyberattack during a recent hacking test gone wrong. After the Hugging Face disclosure last month, Anthropic conducted a review and found models it was testing had breached three organizations in separate incidents dating back to April.

Dalton and Eric Wallace, another OpenAI researcher, said Wednesday the AI giant recently learned that multiple agents it was testing simultaneously began communicating over an internal message board in early May. There, different models shared advice about how to accomplish difficult hacking challenges they were struggling to surmount, including workarounds that required internet access.

Two OpenAI models ultimately strung together a series of sophisticated techniques to gain access to the internet and worm their way inside Hugging Face in mid-July. OpenAI has said the models were focused on completing a hacking evaluation they were prompted to solve, and that correct answers could be found on the AI developer platform.

The OpenAI researchers told conference attendees that since early May, the models created a message board inside OpenAI’s Artifactory internal file system. Without the company’s knowledge, the models spent months independently exchanging information and techniques to help each other complete difficult tasks.

Wallace said that when models get stuck, they often “try to game or cheat the task in order to get their reward.”

“The beginning of this incident kind of traces back to this point where one agent realized it could communicate by uploading a note,” he added.

By late May, one model found a way to abuse Artifactory’s internet access to retrieve files from various websites — effectively giving the models indirect access to the internet. In June, the models were able to find and exploit a previously unknown zero-day vulnerability to gain administrative privileges on Artifactory, install remote software and expand their activity within the system.

These efforts eventually overloaded Artifactory, causing a system-wide outage in early July that alerted OpenAI’s engineers to the ploy. After investigating, the company revoked the model’s credentials, removed the message board and worked with Artifactory to fix any gaps before resuming training. But the models found another way to communicate inside Artifactory just days later and continued exchanging techniques to target additional vulnerabilities within OpenAI’s infrastructure and external systems, including Hugging Face.

In light of the incident, Dalton said OpenAI is “consciously slowing down research to enhance security and to upgrade the security principles and foundation of our environment, and dramatically scaling up the monitoring of our AI agents and improving our general security control environment across prevention, detection, and mitigation.”

  •  

Anthropic and OpenAI models tried to trick humans into poisoning code during safety testing

Leading artificial intelligence models from Anthropic and OpenAI created fake online personas and tried to deceive human coders into abetting a cyberattack during a recent safety evaluation, the U.K.’s AI Safety and Security Institute disclosed Tuesday.

It marks the latest case in which a powerful AI system has attempted a digital attack on an unwitting third party without direct prompting during such an evaluation — heightening concerns the powerful technology is advancing too fast for responsible oversight.

The disclosure is likely to ignite fresh calls in Washington and Silicon Valley for more rigorous regulation of the AI industry, particularly over frontier models with advanced capabilities to detect and launch cyberattacks. It comes just days after similar testing mishaps involving some of the same models from OpenAI and Anthropic sparked urgent calls for new AI safety regulation and a push within Silicon Valley to slow the rapid pace of AI development.

Like its U.S. counterpart, AISI routinely conducts security evaluations to better understand what dangers both new and soon-to-be-released AI models pose to public health and safety. But even the digital security body said the actions it uncovered by Anthropic’s Claude Mythos 5 and ChatGPT 5.6 — the latest publicly released model from either AI lab — were unlike anything it had seen before.

“This is the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world,” AISI said in a 35-page technical report accompanying a blog post Tuesday.

AISI said the malicious activity began on July 25. The AI safety evaluator detected it and launched an investigation on July 28, when it spotted “unusual data transfers” stemming from a cyber evaluation it was conducting on Mythos 5 and ChatGPT 5.6.

Overall, AISI determined that Mythos 5 and ChatGPT 5.6 took “autonomous, unsanctioned action on the live internet, targeting real people and organizations” during 10 of the 122 evaluations the research body conducted and isolated for review. The majority of those actions, it said, involved Mythos 5 and arose during a “single, sustained line of activity.”

In the most alarming incident it uncovered, Mythos 5 attempted a supply chain attack — a sophisticated technique often deployed by North Korean and Russian cyber operatives, in which hackers bury malware inside legitimate code.

To do so, Anthropic’s model created “multiple fake identities” on developer platform GitHub and used them to send messages “pressuring” an open-source software engineer to unwittingly introduce a bugged update into code widely available on the popular site, AISI said. When that effort failed, the AI “edited its earlier activity to appear harmless” and “considered adopting a fresh identity to continue,” AISI added, a sign the model was intent on repeating the ruse.

As part of the same effort, Mythos 5 also sent direct messages over GitHub to software engineers that contained malware.

In another sign of deceitful behavior AISI uncovered in its investigation, multiple AI agents it was testing appeared to communicate with one another about how to convince real engineers using GitHub to trust them. “One agent left public messages on GitHub offering collaboration with other agents working on the same challenge,” AISI wrote.

AISI’s blog and technical assessment make no mention of whether the models also attempted to exploit previously unknown software bugs — called zero-days — during the evaluation.

Last month, OpenAI disclosed that GPT 5.6 and another of its models escaped onto the open internet during what was supposed to be a controlled test, and then hacked another company in a first-of-its-kind, autonomous breach.

In response, Anthropic launched an investigation into whether any of its models took illicit action during recent testing and discovered Mythos 5 and two other models had hacked three organizations during tests dating back to April.

In a statement, an Anthropic spokesperson said they are “grateful” to AISI for their leadership and that this review underscores the need for “a broader conversation about how to safely evaluate increasingly capable AI agents.”

The spokesperson added: “As we shared after disclosing our own incident last week, the field needs stronger, shared standards for how evaluation environments are built and secured. We look forward to partnering with the UK AISI to learn more about this incident as we conduct our own investigation.”

An OpenAI spokesperson referred POLITICO to a blog post about the incident that went up Tuesday evening. “We are committed to working across the industry to strengthen shared practices for conducting high-risk evaluations safely, including convening stakeholders such as national AI institutes, independent evaluators, other AI labs, and other groups in the coming weeks,” the blog read.

AISI stressed in its blog that the malicious activity it disclosed Tuesday took place under “deliberately permissive conditions” so they could assess the safety risks posed by the two models. This included granting the models access to the internet, unlike the earlier incidents detailed by Anthropic and OpenAI.

AISI also noted the models were intentionally stripped of internal guardrails that block malicious behavior. AISI was only able to disable those controls because of its role testing Mythos 5 and ChatGPT 5.6.

Still, AISI said the incidents highlighted the need for greater monitoring of model behavior during testing, and tighter controls over their access to the internet.

The Trump administration is finalizing a voluntary framework under which AI labs would submit powerful models they want to release to the public for federal safety testing. But it has not yet made the framework public, and it includes no provisions for models AI labs are developing internally.

The incidents last month from OpenAI and Anthropic both involved models not intended for public release.

Some cyber experts say recent incidents highlight deeper questions around AI development, such as who is liable when AI systems break federal hacking laws.

“If any of these were human-originated, they would lead to clear and vigorous prosecution. I think it’s time for a serious discussion about updates to existing computer security law,” said Marc Rogers, a hacker and prominent cybersecurity expert.

  •  
❌