“AI Safety” Research Is a Hallucination

‍ ‍Research Funding in “AI Safety” Supports One Conclusion.

That makes it Propaganda - Not Science.

Executive Summary

In recent weeks, a debate about “AI Safety” has exploded into the public eye. This has centered on claims by OpenAI, Anthropic, and affiliated firms and non-profit think tanks about the “existential risk” posed by the advance of Artificial Intelligence. These ideas originate at the intersection of movements concerned with long-term “existential risk,” above all the “Effective Altruist” movement centered around Oxford, and the “Rationalist” movement spearheaded by Eliezer Yudkowsky.

Effective Altruism entails a strain of “utilitarian” thinking that, at its most reductive, can seem to excuse ethical shortcuts in service of a “greater good.” Because they believe themselves to be fighting the total destruction of humanity, even a small chance of having a beneficial impact in the future can excuse an unethical action today.

This report summarizes and expands on reports of research misconduct by elements of the EA/Rationalism network that some scholars argue have operationalized utilitarian thinking into a closed funding system that uncritically perpetuates the movement’s ideas about technology, ethics, and politics. 

We consider this track record important context for the claims being made by Effective Altruists and Rationalists now leading the public discussion about Artificial Intelligence.

This is part one of a two-part report. This installment focuses on academic funding patterns. Part Two will delve into the roles of Coefficient Giving and its smaller counterpart, Survival and Flourishing, in funding journalism about AI.

Subscribe to Zero In

The Coxon Affair

The current surging debate on “AI Safety” may itself have been triggered by events that were less organic than they seemed.

In early September, an Anthropic staffer named Jacob Coxon announced that he was quitting the firm because, he claimed, he believed there was a 10% chance that artificial intelligence “could kill us all by the end of the decade.” Coxon, who had also worked at OpenAI, claimed that “neither company is acting responsibly” about the threat that the rapid advance of machine learning models poses to all of us.

Coxon’s actions have been broadly treated as serious whistleblowing. They came not long after the Hugging Face incident, which OpenAI (and much of the media) characterized as the consequence of agents “going rogue” and acting in unpredictable ways. He told Fox News’ Bret Baier that he had only discussed his decision ahead of time with a few friends, and did not coordinate with third parties in announcing his resignation.

But it quickly became clear that Coxon’s protest was not entirely as it appeared. A Wall Street Journal report on the resignation, including exclusive comment from Coxon, was published 18 minutes before Coxon’s actual announcement. Coxon later said he “misspoke” to Baier and had indeed spoken to the Journal ahead of his announcement. 

Pirate Wires (an outlet funded by anti-Doomer “accelerationist” forces) found that Coxon’s scads of post-resignation interviews were booked partly in collaboration with a public relations firm, DEY.Ideas+Influence, which has also promoted “Doomer in Chief” Eliezer Yudkowsky. Pirate Wires and many of its right-leaning, anti-regulatory cohort have framed the Coxon incident as media manipulation aimed not merely at drawing more attention to safety issues, but at regulatory capture that would benefit Anthropic and OpenAI. 

Just as Coxon announced his doom-driven resignation, many observers were questioning the official narrative about hacks like the Hugging Face incident: OpenAI and Anthropic characterized these as instances of agents acting in uncontrolled ways, but cybersecurity experts have lambasted what they say were negligent security practices.

Discussion of Coxon’s fear of “AI Doom” quickly evolved into a broad public discussion about the profound question of whether a computer intelligence can be “conscious.” Anthropic CEO Dario Amodei said in 2025 that AI models “may be deserving of certain important rights.” In recent days, it has been revealed that Anthropic has courted religious leaders and floated the possibility that its models might be conscious. Most staggering – and likely offensive to many – Anthropic tried to influence the content of Pope Leo XIV’s encyclical “Magnifica Humanitas” to be more open to the possibility of machine consciousness.

This would have the convenient implication both that the development of AI deserves special protections, and that the labs themselves may have less liability in the case of hacks or other harms caused by agents or models. 

More broadly, the “Doom” narrative carries with it the assumption that AI will continue to become rapidly more powerful. Coxon warned that “These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.”

In this sense, the “Doom” narrative, while appearing oppositional, actually serves the interests of the AI labs. This is critical context for the reception of Coxon’s resignation.

Subscribe to Zero In

An Alignment of Interests

Motivated, conflicted, or even dishonest communication is not uncommon, even from supposedly high-minded nonprofits or research organizations. But Effective Altruist and “AI Doom” groups hold explicit “utilitarian” moral views that might, for those within this remarkably tight-knit and proudly heterodox community, justify deceptive communication. 

The underlying utilitarian ethos of Effective Altruism first came to public attention in the case of Sam Bankman-Fried. The disgraced FTX founder’s Effective Altruist/Utilitarian beliefs seemed to justify, in his mind, the theft of billions in customer and investor funds, and to further justify lying about it brazenly both in the media and under oath.

Many other adherents of the movement have demonstrated similarly opaque or even deceptive communication practices. Given the stakes of the current AI debate, this track record deserves serious scrutiny.

This is the first installment of a two-part report surveying two lines of these conflicted influence efforts.

Academic Research Influence:

We have uncovered previously little-known allegations that invite scrutiny of the research ethics of the primary Effective Altruist and “AI Doomer” funding organization, Open Philanthropy (now Coefficient Giving), and of academic figures prominent in the movement.

These figures helped craft a communications agreement under which recipients of research grants from Open Philanthropy were encouraged to cite the research of Open Philanthropy officials and movement leaders, and to shape their work to align with high-level conclusions and priorities of the granting entity. The agreements were allegedly meant to cover any grant recipients who were also employees of universities, possibly violating institutional ethics standards.

We also recount the experience of Carla Cremer, an insider of the “AI Doom” movement who nevertheless experienced serious pushback to her attempt to publish heterodox research, and has alleged serious systematic problems in EA grant funding.

Journalistic Influence: 

Open Philanthropy, along with Survival and Flourishing, are also major funders of journalism fellowships and other money that supports news organizations’ coverage of Artificial Intelligence. In recent days, this funding has become a hot topic, raising questions of the objectivity of the organizations and reporters receiving the funding. Particular attention has focused on the Tarbell Center for AI Journalism, which, in addition to possible ongoing conflicts of interest, has been accused of serious process failures and conflicts of interest in the past.

These questions will be explored in the second part of this report.

Subscribe to Zero In

Why Trusting AI Researchers Matters

Effective Altruism’s most infamous adherent “believed that the ways that people try to justify rules like ‘don’t lie’ [and] ‘don’t steal’, under utilitarianism, didn’t work.”

That, of course, is a description of Sam Bankman-Fried by fellow Effective Altruist Caroline Ellison. Before Bankman-Fried was revealed as a philosophically unrepentant liar, Effective Altruist leaders, including primary figurehead Will MacAskill, themselves downplayed early warnings of his malfeasance.

This example alone justifies suspicion when unpacking the claims of utilitarians leading the EA, Rationalist, and AI Safety communities. It is not implausible that these organizations would base their actions on a moral theory that justifies distorting or exaggerating their messages if they believed the end result would serve their goals.

But examination of the movement’s track record of ethically questionable communication and funding practices suggests something even worse than conscious, active deception. These groups appear to have forgotten or abandoned longstanding methods for ensuring the integrity of research and knowledge claims, including such basics as merit-based funding and openness to conflicting claims. 

Put differently, Effective Altruists and Rationalists have forgotten how to learn rigorous truths, and replaced that with the well-funded promotion of pre-determined conclusions - and the defunding and marginalization of dissenting voices. Major funders and grant allocators have consistently and repeatedly operated on the principle that funding people who already agree with you, or guiding funding in ways that encourage more people to express agreement with you, is not an issue when you are absolutely, positively sure you are right.

This solipsism may seem merely juvenile – the natural ethics of an archipelago of group homes, polycules, and other arrested-development summer camps that the adjacent figure Malcolm Collins recently described as “a giant peerage network who sits with their thumb up their butt all day producing nothing but Harry Potter fan fictions.” 

But an overwhelming proportion of the figures now running and staffing frontier AI labs emerged from this universe of funding, socialization – and, some might argue, indoctrination. If it is a general and accepted practice of this network to manipulate “research” on their core specialties to privilege desired conclusions over open inquiry, that is now a problem with implications for the global economy and the (real, not fantasized) future of the human race.

Whether by design or by emergence from first principles, the result is a system that mimics both intellectual and economic diversity, but dances to a single tune. Just as Sam Bankman-Fried’s Alameda Research posed as an independent entity while effectively acting as an arm of FTX, entities like METR and the Future of Life Institute are, in financial, social, and political terms, clandestinely related third parties of Anthropic and OpenAI.

We can’t call this broader system anything so crude as a conspiracy. There have been discrete conspiracies within it, including some we still don’t understand. But the ‘alignment’ of their funding and communication seems less an active choice, and more a deeply-rooted consequence of their core epistemic principles.

This sense of the rightness of unity can leave many group members without a basic intuition of principles of fairness, transparency, and intellectual rigor. The AI Safety-ists and Altruists seem to believe that by cutting off challenging lines of inquiry, including by defunding or drowning out critics, they are strengthening their own case. 

They see the material strength of their movement as unproblematically vindicating the correctness of their ideas. 

This is an obvious recipe for disaster.

Subscribe to Zero In

The Players

The organizations and individuals who drove the funding patterns described in our reports have, in many cases, transitioned from leadership of Effective Altruist advocacy groups to positions where they fund or lead the “AI Safety” movement. Generally, these are cosmetic changes that reflect a shifting reality, as Artificial Intelligence and its purported “extinction risk” have been the main interest of the Effective Altruist movement for years.

Dustin Moskovitz is a co-founder of Facebook and, later, Asana. His net worth is estimated at nearly $13 billion. With his wife Cari Tuna, Moskovitz co-founded Good Ventures. In 2011, Good Ventures began a partnership with Holden Karnofsky’s Givewell, which eventually became Open Philanthropy. Open Philanthropy, in 2025, rebranded to Coefficient Giving. Coefficient Giving has been the dominant funder of Effective Altruism/Rationalism/AI Safety for most of the movement’s lifespan, only briefly displaced by the FTX Future Fund.

Nick Beckstead is a philosopher whose work focuses on longtermism and the ethics of the far future. He previously served as a program officer at Open Philanthropy and later led the FTX Future Fund's grantmaking. During his tenure as a grant investigator at Open Philanthropy, he wrote communications guidelines that encouraged grant recipients to cite his own research and that of other fund administrators. Like many Effective Altruists, he has since turned his interest to AI Safety, as CEO of the Secure AI Project.

Holden Karnofsky co-founded GiveWell and Open Philanthropy, where he has helped direct large-scale giving toward global health and catastrophic-risk reduction. He is the former CEO of Open Philanthropy (Coefficient Giving), and served on the OpenAI board from 2017 to 2021.

Will MacAskill is a moral philosopher at Oxford and a co-founder of the effective altruism movement, including Giving What We Can and 80,000 Hours. He is the author of Doing Good Better, What We Owe the Future, and other books on effective altruism and longtermism. In his role as an advisor to the FTX Future Fund, MacAskill oversaw the distribution of roughly $160 million in grants consisting substantially of stolen customer funds. About $36 million of that money went to organizations that MacAskill also led or helped found.

While touting his commitment to donating a large portion of his income, MacAskill appears to have used millions of dollars of Effective Altruist donor funds to promote his own career. In 2023, in the wake of the FTX collapse, MacAskill resigned from Effective Ventures UK.

Toby Ord is also a moral philosopher at Oxford who co-founded Giving What We Can and helped launch the effective altruism movement. He is the author of The Precipice, which examines existential risks to humanity.

Beckstead, Karnofsky, MacAskill, and Ord were all listed as having reviewed the Open Philanthropy version of the communication agreement described below.

Coefficient Giving (Formerly Open Philanthropy)

Primarily but not exclusively a funnel for the wealth of Dustin Moskovitz, Coefficient Giving has constituted a huge proportion of all Effective Altruist funding throughout the movement’s history. Numbers compiled circa 2019 implied 62.5% of all EA funding was from Open Philanthropy at that time. The appearance of FTX in 2022 briefly changed the mix, but an estimate based on mid-2023 data suggests OpenPhil quickly regained its dominant funding position.

This dominance is important because many of the dubiously ethical communication practices described below can be traced back to Open Philanthropy/Coefficient Giving. Among many, many other projects, Open Philanthropy is a funder of the Tarbell Center for AI Journalism. While it does distribute funds from other sources, Moskovitz remains its primary funder by a substantial margin.

The organization has given on the order of $4 billion to EA and AI Safety organizations in total, creating a vast network of seemingly independent entities actually beholden, in practice, to one man.

Can you Trust a Utilitarian?

The events below include behavior that would, under the stewardship of a normal business or nonprofit, be so obviously misguided as to invite serious questions about the players’ state of mind. This strangeness is rooted in the “utilitarian” ethics shared by many in the AI Safety, Rationalist, and Effective Altruism movements.

Utilitarianism broadly values general ethical principles less than decisions based on an action’s overall net benefit. In the cases of conflict of interest, information control, and strategic funding below, we see a milder but perhaps more insidious utilitarianism that in practice treats intellectual conflict, and public disagreement most of all, as having a negative “expected value.”

Organizations and individuals have varying relationships to Effective Altruism both as an ideology and a set of organizations. But very close to the center of it all sits Coefficient Giving, the single largest explicitly “Effective Altruist” granting organization. It grew out of a partnership between Good Ventures and GiveWell, a charity evaluator cofounded by Holden Karnofsky, an avowed pillar of EA who also figures in events recounted below.

Effective Altruists and other AI “Doomers” have a long track record of using their vast resources to aggressively shape public narratives. Funds disbursed as grants through a network of closely linked nonprofits have, for going on a decade, fostered an ecosystem of bloggers, journalists, and even scholars writing about existential risk and technological progress.

This gives the appearance of a normal nonprofit funding ecosystem, but looking closer reveals a pattern of concentrated funding sources, concealed agreements, conflicts of interest – and, in some cases, intimidation and deception. Some would argue these nonprofits are largely funding not research, but propaganda aimed at swaying public opinion in directions that benefit frontier AI labs. 

Subscribe to Zero In

EA Foundation Communication Agreements 

Some of what we recount here has been reported previously. But one particularly damning episode of ethically questionable funding practices in Effective Altruism and AI Safety has received little public attention because it unfolded in the world of academia.

At issue is a 2019 grant of $1 million to the Effective Altruism Foundation (EAF) from Moskodvitz and Tuna’s Open Philanthropy Project. The grant was tiny by contemporary standards, but according to two whistleblowers, it came with strings attached. Grant recipients were allegedly pressured to comply with these agreements to secure future funding, pointing to deep systemic problems with the EA funding ecosystem. Knutsson has preserved copies of both the EAF version and the Open Philanthropy/Beckstead version of the agreement, as of mid-2019.

Despite its all-encompassing name, the Germany-based EAF represented a somewhat heterodox wing of the EA movement. It was more open than most “extinction risk” or “longtermist” think-tanks to the idea of a “pessimistic” future. That is, their work entertained the possibility that a future civilization with a greater human population might actually increase the total of human suffering, if that civilization exploited or abused enough of its members. This concern with suffering is a major challenge to utilitarianism as such, closely related to Derek Parfit’s “repugnant conclusion.” 

“Pessimistic” risk studies would specifically oppose the utilitarian viewpoint, widespread at OpenAI and Anthropic, that present humans should make sacrifices to accelerate technological development. This viewpoint was recently, eerily echoed by OpenAI CEO Sam Altman, who told Politico on October 4 that “We believe the world should accept some bad things happening for the benefits of this technology.”

An unflinchingly optimistic take on the expanding human population is very useful for those seeking present sacrifices to accelerate the development of artificial intelligence. It appears Coefficient Giving/Open Philanthropy has, at least in the past, used strategic funding to suppress pessimistic ideas that might undermine this case for “accepting some bad things.”

Simon Knutsson was one of the scholars offered funding by the Effective Altruism Foundation from the Open Philanthropy grant. Discussing this prospective funding in June 2019, he stated he was shown two “Communication Guidelines” documents drafted in conjunction with the grant. He was "told essentially that I could likely get funding in the future if I would follow the guidelines."

The two documents broadly discouraged researchers who received EAF grants from focusing on the possibility that the future could be worse than the present, which reversed the organization’s true stance. “In general,” one of the documents reads, “we recommend writing about practical ways to reduce s-risk [suffering risk] without mentioning how the future could be bad overall."

On its face, this is evidence of serious academic research misconduct. Intellectual independence from funding sources was a foundational pillar of Vannevar Bush’s 1945 treatise Science, the Endless Frontier—a document that laid the groundwork for American academic institutions and nearly a century of scientific and technological dominance.

But the Effective Altruists, not for the last time, thought they knew better.

One of the two documents was attributed to Open Philanthropy and attributed to then-grant investigator Nick Beckstead. The version of Beckstead’s guidelines document preserved by Knutsson says it was “endorsed by the following organizations and individuals… 80,000 Hours, CEA, CFAR, MIRI, Open Phil, Nick Bostrom, Will MacAskill, Toby Ord, Carl Shulman.” 

Bostrom, Ord and MacAskill’s presence on this document is particularly problematic, as all three hold positions at Oxford, whose institutional ethics guidelines they may have violated by participating. We have contacted Oxford regarding those policies and will update this story with any response.

The result of this financially enforced bias was the downplaying of the risk of negative future outcomes, including negative outcomes of technological innovation, by the previously “pessimist” Effective Altruist Foundation. 

Torges wrote to Knutsson that blunting this part of the organization’s agenda was a strategic tradeoff:

"We’re excited about this coordinated effort… Even though it means that we will promote future pessimism and s-risks [suffering risk] to a lesser extent, we think the discourse about these topics will still be improved due to other more widely-read texts taking our perspective into account more.”

Crucially, Knutsson understood from Torges that these guidelines are not only “binding for those employed or funded by EAF/FRI, but they are also meant for independent researchers and university employees.” 

This would appear to place the funding organizers and recipients who accepted the guidelines in conflict with the ethics provisions of most major universities: Oxford, where Bostrom, MacAskill and Ord remain employed, has ethical guarantees of Academic Freedom (section 2.2), bars on “suppression of inconvenient results” (4.2), and requires the management of faculty conflicts of interest (3.2, 4.2). (We reached out to Oxford for any comment on their ethics policy, and will update this story if and when we receive a reply.)

Knutsson ultimately declined to accept funding from EAF, or to agree to the communications guidelines. Nonetheless, he says Torges attempted to pressure him towards compliance with the “optimistic” conclusions outlined in the document. This included Torges pressuring Knutsson to change the title of a paper that had already been accepted for publication.

Though Torges was the point of contact for what Knutsson experienced as intellectual bullying, Knutsson is clear he regards Torges more as a misguided functionary than a villain. “My impression is that he is a nice guy, wants to help others, especially those who suffer, and he thinks that this is the best way to do so.” 

The same good intentions could be attributed to many in the Effective Altruism movement. But here, as in many other cases, those intentions, filtered through the strange utilitarian ethics of EA, often manifest in behavior that non-EAs may see as bizarre, counterproductive, or even straightforwardly unethical.

Subscribe to Zero In

A second source, Magnus Vinding, also saw the coordination documents and subsequently departed the mainstream Effective Altruism movement to found the independent Center for Reducing Suffering, where Simon Knutsson now does part-time editorial work. More than the intellectual dishonesty of the coordination effort, Vinding was offended by the push to elevate perspectives that excuse present suffering in exchange for a supposed future bounty.

“Many of the influential EAs such as MacAskill and Ord,” Vinding says, “Endorse axiologies or theories of value that entail that it is worth creating extreme suffering for some in order to boost happiness for the already happy.”

Knutsson’s assessment, made nearly seven years ago, was that “One can also accuse [EA funders] of contributing to what one might call far-future fanaticism, according to which what happens nowadays, even violence and suffering, is essentially negligible as long as the far-future is very good.”

Altman’s invocation of this logic of sacrifice is particularly notable because he is only tangentially associated with Effective Altruism and “extinction risk” research. In fact, he was briefly ousted from OpenAI by the EA wing of that organization’s leadership. Yet he appears to share their consequentialist utilitarian ethics. This demonstrates just how profoundly and subtly the application of research and communications funding towards a crass utilitarian logic has impacted the thinking, not just of a small group of philosophers, but of a huge cadre of fellow-travelers orbiting OpenAI and Anthropic.

The EAF document does specifically state that “we do not ask you to lie or distort your views… Whenever there are conflicts between cooperation and honesty that you can’t resolve, we would like you to side with honesty.”

But this is a deeply specious denial: As I’ve written so many times about statements by Bankman-Fried, it contains an accidental confession.

If your funding guidelines must insist that they are not asking you to lie, you’re admitting that the very existence of the guidelines applies pressure for scholars to bend the curve of their conclusions.  You are describing a situation in which a scholar would have to make a decision between serving their granting organization and serving the truth. No such decision should ever face a researcher.

It is also a frequent topic of discussion within Effective Altruism and AI Safety circles that unorthodox or contrary research can lead to the withdrawal of funding or the loss of future funding or employment opportunities within the ecosystem.

Knutsson believes, very reasonably, that he might have qualified for further funding or even employment if he had “played along” with the guidelines propagated by EA organizations. Instead, he has largely exited that movement and is finishing a more conventional academic PhD path at Stockholm University. 

Instead of collecting grants, he says he picks up work as a carpenter and warehouse worker to make ends meet.

Subscribe to Zero In

Carla Cremer & the Case for Democracy

A second, very similar account of distorting pressure on the research process comes from Carla Zoe Cremer, a former researcher at Nick Bostrom’s own Future of Humanity Institute. Despite this institutional perch within the EA network, Cremer in 2021 described facing massive internal and community resistance to a research project aimed at improving the research practices of Existential Risk researchers.

Cremer at the time described a harrowing research process that, to her, indicated deep intellectual dysfunction. “The creation of this paper has not signalled epistemic health,” she wrote after its publication. “It has been the most emotionally draining paper we have ever written... the burden of proof placed on our claims was unbelievably high in comparison to papers which were considered less ‘political’ or simply closer to orthodox views. Making the case for democracy was heavily contested, despite reams of supporting empirical and theoretical evidence.” (emphasis added)

Specifically, Cremer’s paper critiqued the “techno-utopian approach” prevalent in the MacAskill/Bostrom/Ord wing of Existential Risk scholarship, arguing that it “relies on a non-representative moral worldview, uses ambiguous and inadequate definitions, fails to incorporate insights from risk assessment in relevant fields, chooses arbitrary categorisations of risk, and advocates for dangerous mitigation strategies. Its moral and empirical assumptions might be particularly vulnerable to securitisation and misuse.”

In a spectacular display of the movement’s blindness to its own failings, Cremer claimed that “we were told by some that our critique is invalid because the community is already very cognitively diverse and in fact welcomes criticism.” That absurdity, where one claims that criticism should be silenced because its targets are so open to criticism, is just a small illustration of how the bias of agreement fostered by selective funding produces bad, self-serving thinking at scale.

Cremer found the network’s selective funding bias was already widely accepted by researchers in the network. Figures including “senior scholars in the field… told us in private that they were concerned that any critique of central figures in EA would result in an inability to secure funding from EA sources, such as OpenPhilanthropy." (Emphasis added).

We don't know if these concerns are warranted. Nonetheless, any field that operates under such a chilling effect is neither free nor fair. Having a handful of wealthy donors and their advisors dictate the evolution of an entire field is bad epistemics at best and corruption at worst.

This funding bias operationalizes a screaming flaw in how this movement’s leaders and funding organizations conceive of knowledge: as something that is weakened, rather than strengthened, by dissent. The constant implicit assumption is that funding people who already agree with you, or guiding funding in ways that encourage more people to express agreement with you, is not an issue when you are absolutely, positively sure you are right.

Cremer’s efforts to reform the study of extinction risk did not benefit their standing within Effective Altruism. Their appointment at the Future of Humanity Institute ended in 2021; the institute itself was permanently shuttered by Oxford in 2024, largely due to the severe reputational damage associated with the FTX fraud. Cremer later moved on to Chris Summerfield’s Human Information Processing Lab at Oxford, which has fewer apparent ties to the EA wing of AI Safety, and is headed by a real computer scientist. Cremer has since moved on from HIPL, and has not posted in the Effective Altruism forum under her own name for three years. Attempts to reach her through a listed email address were unsuccessful.

The Dangers of “Alignment”

These examples show “aligned” researchers in the Effective Altruist and AI Safety universe advancing shared, agreed-upon conclusions, not merely through their published arguments, but through back-channel pressure campaigns and strategic funding. These groups appear to believe that such unity strengthens their movement, and from a short-term institutional viewpoint, that may be true. 

But it has become increasingly clear that the cost of such enforced uniformity is the elevation of less competent researchers on the basis of loyalty and agreement, rather than talent and rigor. Telegraph columnist Andrew Orlowski has been tracking the growing influence of Effective Altruism in UK politics, and argues that “What’s called ‘AI Safety’ is a completely made-up and tightly policed field, owned entirely by the EA cult. The... jargon of ‘alignment’ and ‘evals’ etc., is not shared by anyone outside the cult.” 

This “cult” now includes people responsible for the hands-on work of policing AI agent behavior. But cybersecurity professionals speaking to Orlowski described incident reports from OpenAI as “technically incoherent,” and much of the labs’ cybersecurity work is outsourced. This casts systematic doubt on the credibility of their claims that, for instance, agents are “going rogue” and acting unpredictably, rather than simply being poorly designed and irresponsibly supervised. 

The use of money to shape research outputs is hardly novel. We have seen it before, undertaken by the tobacco industry, the fossil fuel industry, the sugar industry, and social media firms. 

Those efforts often hid clandestine research funding by moving it through nonprofit organizations with opaque names like “The Advancement of Sound Science Coalition” and the “Global Climate Coalition.” The campaigns not only enabled harmful practices, but also eroded the general principles of neutrality and truth-seeking at the heart of American economic and industrial dominance.


It may be happening again.