We interview political appointees and civil servants about how policy actually gets made.
Subscribe at www.statecraft.pub to get interview transcripts in your inbox once a week.
Breakneck is like those letters: it goes all over the place, as does our conversation. Topics include:
* America's overabundance of lawyers
* Whether our ruling class should be all economists
* Stylish propaganda
* The book collections of Yale professors
* iPhone manufacturing
* Forced sterilization
* Planting cassava
One of the things I like most about Dan's work is that he's comfortable looking at China through multiple, very different lenses. Parts of Breakneck explicitly use China as a lens to think about the US and its political culture and institutions. Other parts of the book try very hard to take China on its own terms, without reading our own culture into it. It’s that mix that made the book so enjoyable for me, and I hope you enjoy it too.
Thank you to Harry Fletcher-Wood for his judicious transcript edits, and to Katerina Barton for her audio edits. You can find the full, annotated transcript to this conversation at www.statecraft.pub.
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.statecraft.pub
Four Ways to Fix Government HR
jeudi 21 août 2025 • Durée 01:03:02
Today I'm talking to economic historian Judge Glock, Director of Research at the Manhattan Institute. Judge works on a lot of topics: if you enjoy this episode, I'd encourage you to read some of his work on housing markets and the Environmental Protection Agency. But I cornered him today to talk about civil service reform.
Since the 1990s, over 20 red and blue states have made radical changes to how they hire and fire government employees — changes that would be completely outside the Overton window at the federal level. A paper by Judge and Renu Mukherjee lists four reforms made by states like Texas, Florida, and Georgia:
* At-will employment for state workers
* The elimination of collective bargaining agreements
* Giving managers much more discretion to hire
* Giving managers much more discretion in how they pay employees
Judge finds decent evidence that the reforms have improved the effectiveness of state governments, and little evidence of the politicization that federal reformers fear. Meanwhile, in Washington, managers can’t see applicants’ resumes, keyword searches determine who gets hired, and firing a bad performer can take years. But almost none of these ideas are on the table in Washington.
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit
How to Be a Good Intelligence Analyst
jeudi 7 août 2025 • Durée 01:01:26
Today we're joined by Dr. Rob Johnston. He's an anthropologist, an intelligence community veteran, and author of the cult classic Analytic Culture in the US Intelligence Community, a book so influential that it's required reading at DARPA. But first and foremost, Johnston is an ethnographer. His focus in that book is on how analysts actually produce intelligence analysis.
Johnston answers a lot of questions I've had for a while about intelligence and spying, such as:
* Why do we seem to get big predictions wrong so consistently?
* Why can't the CIA find analysts who speak the language of the country they're analyzing?
* Why do we prioritize expensive satellites over human intelligence?
We also discuss a meta-question I always come back to on Statecraft: is being good at this stuff an art or a science? By “this stuff,” I’m referring to intelligence analysis, but I think that the question generalizes across policymaking. Would more formalizing and systematizing make our spies, diplomats, and EPA bureaucrats better? Or would it lead to more bureaucracy, more paper, and worse outcomes? How do you build processes in the government that actually make you better at your job?
You can find the full transcript for this conversation at www.statecraft.pub.
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.statecraft.pub
How to Fix Foreign Aid
jeudi 31 juillet 2025 • Durée 01:14:01
We’ve covered the US Agency for International Development, or USAID, pretty consistently on Statecraft, since our first interview on PEPFAR, the flagship anti-AIDS program, in 2023. When DOGE came to USAID, I was extremely critical of the cuts to lifesaving aid, and the abrupt, pointlessly harmful ways in which they were enacted. In March, I wrote, “The DOGE team has axed the most effective and efficient programs at USAID, and forced out the chief economist, who was brought in to oversee a more aggressive push toward efficiency.”
Today, we’re talking to that forced-out chief economist, Dean Karlan. Dean spent two and a half years at the helm of the first-ever Office of the Chief Economist at USAID. In that role, he tried to help USAID get better value from its foreign aid spending. His office shifted $1.7 billion of spending towards programs with stronger evidence of effectiveness. He explains how he achieved this, building a start-up within a massive bureaucracy. I should note that Dean is one of the titans of development economics, leading some of the most important initiatives in the field (I won’t list them, but see herefor details), and I think there’s a plausible case he deserves a Nobel.
Throughout this conversation, Dean makes a point much better than I could: the status quo at USAID needed a lot of improvement. The same political mechanisms that get foreign aid funded by Congress also created major vulnerabilities for foreign aid, vulnerabilities that DOGE seized on. Dean believes foreign aid is hugely valuable, a good thing for us to spend our time, money, and resources on. But there's a lot USAID could do differently to make its marginal dollar spent more efficient.
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit
How Cheaply Could We Build High-Speed Rail?
mercredi 23 juillet 2025 • Durée 56:32
At the end of April, the Transit Costs Project released a report: it’s called How to Build High-Speed Rail on the Northeast Corridor. As the name suggests, the authors of the report had a simple goal: the stretch of the US from DC and Baltimore through Philadelphia to New York and up to Boston, the densest stretch of the country. It’s an ideal location for high-speed rail. How could you actually build it — trains that get you from DC to NYC in two hours, or NYC to Boston in two hours — without breaking the bank?
That last part is pretty important. The authors think you could do it for under $20 billion dollars. That’s a lot of money, but it’s about five times less than the budget Amtrak says it would require. What’s the difference? How is it that when Amtrak gets asked to price out high-speed rail, it gives a quote that much higher?
We brought in Alon Levy, transit guru and the lead author of the report, to answer the question, and to explain a bunch of transit facts to a layman like me. Is this project actually technically feasible? And, if it is, could it actually work politically?
* How to cut time off the Northeast Corridor
* Operations coordination as a time-saver
* The move away from the Mad Men commuter
* Was our episode on the Green Line extension wrong?
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit
Governance Lessons From the Constitutional Convention
vendredi 4 juillet 2025 • Durée 13:42
Happy Fourth of July! I’m attending a wedding today, so this episode is from the vault, in a way, although it’s its first time on Statecraft. I originally published this essay in January of 2022 on Mirror, shortly after my wife had joined the core team of a DAO that was attempting to acquire a first-edition copy of the US Constitution. I had been reading a history of the constitutional convention, and it seemed fitting to write about it on a thematic site. Yes, July 4th is about the Declaration of Independence, not the Constitution. Cut me some slack, please!You can find the transcript for this episode and many others at www.statecraft.pub.
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.statecraft.pub
How to Predict the Future
mercredi 25 juin 2025 • Durée 01:24:18
The decisions that humans make can be extraordinarily costly. The wars in Iraq and Afghanistan were multi-trillion-dollar decisions. If you can improve the accuracy of forecasting individual strategies by just a percentage point, that would be worth tens of billions of dollars. Yet society does not invest tens of billions of dollars in figuring out how to improve the accuracy of human judgment. That seems really odd.
That’s a quote from today’s interviewee, who has made his career helping the intelligence community predict the future better. In this interview, we discuss:
* Which prediction methods perform the best?
* How does IARPA create tech for American spies?
* What technologies give democracies an advantage over autocracies?
* Could the Internet have been designed better?
Our interviewee, Jason Matheny, championed research into human judgment and forecasting at the R&D lab for the intelligence community: the Intelligence Advanced Research Projects Activity, or IARPA, which he directed from 2015-2018.
[This interview was originally published in 2023, at this link, without the audio: Statecraft was still transcript-only then.]
You can find the transcript for this conversation at www.statecraft.pub.
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.statecraft.pub
How UK Biobank Was Built
jeudi 19 juin 2025 • Durée 48:45
There are many forces in policymaking (and in our lives generally) that push us towards the short term. Many of the most important measurements in political life are on extremely tight timelines: election cycles, monthly unemployment reports, even the President's daily intelligence briefing. The pressure to get results — and to show results — on a tight turnaround is incredible.
One of my questions on Statecraft for a while has been: How do you build a machine to get long-term results? Whether it's a new agency or a new initiative, how do you set up a structure to work toward a goal that's 10, or 20, or 50 years away? And how do you protect that structure from short-term political pressures?
Today's interviewee is Sir Rory Collins. Sir Rory has spent a full 20 years building and leading one of the most important scientific resources in the world: the UK Biobank.
The Biobank represents a fascinating case study in long-term thinking. It's a database of half a million British participants whose health is being tracked longitudinally for the next 30 years. The Biobank was established with the knowledge that the upfront work, and the spending required, would only really start to pay off 15 years later. When Sir Rory went in for the 10-year review with funders, they asked what had been achieved so far. He said, “Nothing.”
But today, UK Biobank is paying massive dividends: It's democratized access to population-scale data for researchers worldwide, and it's already yielding amazing insights into the causes of and cures for disease. I wanted to understand how he built the UK Biobank, and, just as importantly, how he managed to sustain it over a long period of time.
We discussed
* How to create long-term value in research
* How to recruit half a million research subjects
* Why the Biobank deferred so many decisions
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit
What Can We Learn From Estonia?
jeudi 12 juin 2025 • Durée 46:18
What can we learn from Estonia? It’s not a question you hear often — the nation of under two million residents doesn’t mean much to many. But for good governance advocates, it’s long been a touchpoint for its “e-government” model. The New Yorker wrote in 2017 that, “apart from transfers of physical property, such as buying a house, all bureaucratic processes can be done online.” Wired called Estonia “the world's most digitally advanced society.” On its “e-Estonia” site, the country itself brags, in a mod font, “We have built a digital society and we can show you how.”
The Estonian model has a lot going for it from the perspective of a citizen. For example: Taxes take a few minutes to file, you can see every time the government looks at your data, and you never have to give the government a piece of information more than once. And it makes governance easier: the bureaucracy is leaner, information is shared across agencies, and data is more secure.
But how much of this model could be adopted here in the US, or in the rest of the West? And how much is reliant on a cultural and societal context we just don’t have here? To get answers, I talked to Joel Burke, author of the new book Rebooting a Nation: The Incredible Rise of Estonia, E-Government and the Startup Revolution. Joel is an American who worked with the Estonian government, and I learned a lot from his book.
For the full transcript of this conversation and others, visit www.statecraft.pub.
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit
How to Save DC's Metro
jeudi 5 juin 2025 • Durée 33:13
Today we talked to Randy Clarke, the head of DC’s Metro system, WMATA. If you’re a transit nerd, you probably know about Clarke — he’s become something of a celebrity for his public presence and disciplined improvement of a transit system that was facing disaster in the aftermath of COVID (and the decision to allow large swathes of federal employees to work from home).
I’ve been a regular WMATA rider for long periods of my life, and what Clarke has done over the last three years has been pretty remarkable. We’ll get into some of the details here, but what stands out to me — and why I so wanted to record this conversation — is that Clarke’s managed to advance a bunch of his priorities at once. From the outside, it can seem like he hasn’t had to make any tradeoffs at all: between safety and speed, catching fare evaders and keeping costs down, etc. How has he pulled it off?
You can read the transcript for this conversation (and others) at www.statecraft.pub.
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.statecraft.pub
Thanks to Harry Fletcher-Wood for his judicious transcript edits and fact-checking, and to Katerina Barton for audio edits.
Judge, you have a paper out about lessons for civil service reform from the states. Since the ‘90s, red and blue states have made big changes to how they hire and fire people. Walk through those changes for me.
I was born and grew up in Washington DC, heard a lot about civil service throughout my childhood, and began to research it as an adult. But I knew almost nothing about the state civil service systems. When I began working in the states — mainly across the Sunbelt, including in Texas, Kansas, Arizona — I was surprised to learn that their civil service systems were reformed to an absolutely radical extent relative to anything proposed at the federal level, let alone implemented.
Starting in the 1990s, several states went to complete at-will employment. That means there were no official civil service protections for any state employees. Some managers were authorized to hire people off the street, just like you could in the private sector. A manager meets someone in a coffee shop, they say, "I'm looking for exactly your role. Why don't you come on board?" At the federal level, with its stultified hiring process, it seemed absurd to even suggest something like that.
You had states that got rid of any collective bargaining agreements with their public employee unions. You also had states that did a lot more broadbanding[creating wider pay bands] for employee pay: a lot more discretion for managers to reward or penalize their employees depending on their performance.
These major reforms in these states were, from the perspective of DC, incredibly radical. Literally nobody at the federal level proposes anything approximating what has been in place for decades in the states. That should be more commonly known, and should infiltrate the debate on civil service reform in DC.
Even though the evidence is not absolutely airtight, on the whole these reforms have been positive. A lot of the evidence is surveys asking managers and operators in these states how they think it works. They've generally been positive. We know these states operate pretty well: Places like Texas, Florida, and Arizona rank well on state capacity metrics in terms of cost of government, time for permitting, and other issues.
Finally, to me the most surprising thing is the dog that didn't bark. The argument in the federal government against civil service reform is, “If you do this, we will open up the gates of hell and return to the 19th-century patronage system, where spoilsmen come and go depending on elected officials, and the government is overrun with political appointees who don't care about the civil service.” That has simply not happened. We have very few reports of any concrete examples of politicization at the state level. In surveys, state employees and managers can almost never remember any example of political preferences influencing hiring or firing.
One of the surveys you cited asked, “Can you think of a time someone said that they thought that the political preferences were a factor in civil service hiring?” and it was something like 5%.
It was in that 5-10% range. I don't think you'd find a dissimilar number of people who would say that even in an official civil service system. Politics is not completely excluded even from a formal civil service system.
A few weeks ago, you and I talked to our mutual friend, Don Moynihan, who's a scholar of public administration. He's more skeptical about the evidence that civil service reform would be positive at the federal level.
One of your points is, “We don't have strong negative evidence from the states. Productivity didn’t crater in states that moved to an at-will employment system.” We do have strong evidence that collective bargaining in the public sector is bad for productivity.
What I think you and Don would agree on is that we could use more evidence on the hiring and firing side than the surveys that we have. Is that a fair assessment?
Yes, I think that's correct. As you mentioned, the evidence on collective bargaining is pretty close to universal: it raises costs, reduces the efficiency of government, and has few to no positive upsides.
On hiring and firing, I mentioned a few studies. There's a 2013 study that looks at HR managers in six states and finds very little evidence of politicization, and managers generally prefer the new system. There was a dissertation that surveyed several employees and managers in civil service reform and non-reform states. Across the board, the at-will employment states said they had better hiring retention, productivity, and so forth. And there's a2002 study that looked specifically at Texas, Florida, and Georgia after their reforms, and found almost universal approbation inside the civil service itself for these reforms.
These are not randomized control trials. But I think that generally positive evidence should point us directionally where we should go on civil service reform. If we loosen restrictions on discipline and firing, decentralize hiring and so forth — we probably get some productivity benefits from it. We can also know, with some amount of confidence, that the sky is not going to fall, which I think is a very important baseline assumption. The civil service system will continue on and probably be fairly close to what it is today, in terms of its political influence, if you have decentralized hiring and at-will employment.
As you point out, a lot of these reforms that have happened in 20-odd states since the ‘90s would be totally outside the Overton window at the federal level. Why is it so easy for Georgia to make a bipartisan move in the ‘90s to at-will employment, when you couldn't raise the topic at the federal level?
It's a good question. I think in the 1990s, a lot of people thought a combination of the 1978 Civil Service Reform Act — which was the Carter-era act that somewhat attempted to do what these states hoped to do in the 1990s — and the Clinton-eraReinventing Government Initiative, would accomplish the same ends. That didn't happen.
That was an era when civil service reform was much more bipartisan. In Georgia, it was a Democratic governor, Zell Miller, who pushed it. In a lot of these other states, they got buy-in from both sides. The recent era of state reform took place after the 2010 Republican wave in the states. Since that wave, the reform impetus for civil service has been much more Republican. That has meant it's been a lot harder to get buy-in from both sides at the federal level, which will be necessary to overcome a filibuster.
I think people know it has to be very bipartisan. We're just past the point, at least at the moment, where it can be bipartisan at the federal level. But there are areas where there's a fair amount of overlap between the two sides on what needs to happen, at least in the upper reaches of the civil service.
It was interesting to me just how bipartisan civil service reform has been at various times. You talked about the Civil Service Reform Act, which passed Congress in 1978. President Carter tells Congress that the civil service system:
“Has become a bureaucratic maze which neglects merit, tolerates poor performance, permits abuse of legitimate employee rights, and mires every personnel action in red tape, delay, and confusion.”
That's a Democratic president saying that. It’s striking to me that the civil service was not the polarized topic that it is today.
Absolutely. Carter was a big civil service reformer in Georgia before those even larger 1990s reforms. He campaigned on civil service reform and thought it was essential to the success of his presidency. But I think you are seeing little sprouts of potential bipartisanship today, like the Chance to Compete Act at the end of 2024, and some of the reforms Obama did to the hiring process. There's options for bipartisanship at the federal level, even if it can’t approach what the states have done.
I want to walk through the federal hiring process. Let's say you're looking to hire in some federal agency — you pick the agency — and I graduated college recently, and I want to go into the civil service. Tell me about trying to hire somebody like me. What's your first step?
It's interesting you bring up the college graduate, because that is one recent reform: President Trump put out an executive order trying to counsel agencies to remove the college degree requirement for job postings. This happened in a lot of states first, like Maryland, and that's also been bipartisan. This requirement for a college degree — which was used as a very unfortunate proxy for ability at a lot of these jobs — is now being removed. It's not across the whole federal government. There's still job postings that require higher education degrees, but that's something that's changed.
To your question, let’s say the Department of Transportation. That's one of the more bipartisan ones, when you look at surveys of federal civil servants. Department of Defense, Veterans Affairs, they tend to be a little more Republican. Health and Human Services and some other agencies tend to be pretty Democrat. Transportation is somewhere in the middle.
As a manager, you try to craft a job description and posting to go up on the USA Jobs website, which is where all federal job postings go. When they created it back in 1996, that was supposedly a massive reform to federal hiring: this website where people could submit their resumes. Then, people submit their resumes and answer questions about their qualifications for the job.
One of the slightly different aspects from the private sector is that those applications usually go to an HR specialist first. The specialist reviews everything and starts to rank people into different categories, based on a lot of weird things. It's supposed to be “knowledge, skills, and abilities” — your KSAs, or competencies. To some extent, this is a big step up from historical practice. You had, frankly, an absurd civil service exam, where people had to fill out questions about, say, General Grant or about US Code Title 42, or whatever it was, and then submit it. Someone rated the civil service exam, and then the top three test-takers were eligible for the job.
We have this newer, better system, where we rank on knowledge, skills, and abilities, and HR puts put people into different categories. One of the awkward ways they do this is by merely scanning the resumes and applications for keywords. If it's a computer job, make sure you say the word “computer” somewhere in your resume. Make sure you say “manager” if it's a managerial job.
Just to be clear, this is entirely literal. There's a keyword search, and folks who don’t pass that search are dinged.
Yes. I've always wondered, how common is this? It's sometimes hard to know what happens in the black box in these federal HR departments. I saw an HR official recently say, "If I'm not allowed to do keyword searches, I'm going to take 15 years to overlook all the applications, so I’ve got to do keyword searches." If they don't have the keywords, into the circular file it goes, as they used to say: into the garbage can.
Then they start ranking people on their abilities into, often, three different categories. That is also very literal. If you put in the little word bubble, "I am an exceptional manager," you get pushed on into the next level of the competition. If you say, "I'm pretty good, but I'm not the best," into the circular file you go.
I’ve gotten jaded about this, but it really is shocking. We ask candidates for a self-assessment, and if they just rank themselves 10/10 on everything, no matter how ludicrous, that improves their odds of being hired.
That's going to immensely improve your odds. Similar to the keyword search, there's been pushback on this in recent years, and I'm definitely not going to say it's universal anymore. It's rarer than it used to be. But it’s still a very common process.
The historical civil service system used to operate on a rule of three. In places like New York, it still operates like that. The top three candidates on the evaluation system get presented to the manager, and the manager has to approve one of them for the position.
Thanks partially to reforms by the Obama administration in 2010, they have this category rating system where the best qualified or the very qualified get put into a big bucket together [instead of only including the top three]. Those are the people that the person doing the hiring gets to see, evaluate, and decide who he wants to hire.
There are some restrictions on that. If a veteran outranks everybody else, you’ve got to pick the veteran [typically known as Veterans’ Preference]. That was an issue in some of the state civil service reforms, too. The states said, “We're just going to encourage a veterans’ preference. We don't need a formalized system to say they get X number of points and have to be in Y category. We're just going to say, ‘Try to hire veterans.’” That’s possible without the formal system, despite what some opponents of reform may claim.
One of the particular problems here is just the nature of the people doing the hiring. Sometimes you just need good managers to encourage HR departments to look at a broader set of qualifications. But one of the bigger problems is that they keep the HR evaluation system divorced from the manager who is doing the hiring. David Shulkin, who was the head of the Department of Veterans Affairs (VA), wrote a great book, It Shouldn’t Be This Hard to Serve Your Country. He was a healthcare exec, and the VA is mainly a healthcare agency. He would tell people, "You should work for me," they would send their applications into the HR void, and he'd never see them again. They would get blocked at some point in this HR evaluation process, and he'd be sent people with no healthcare experience, because for whatever reason they did well in the ranking.
One of the very base-level reforms should be, “How can we more clearly integrate the hiring manager with the evaluation process?” To some extent, the bipartisan Chance to Compete Act tries to do this. They said, “You should have subject matter experts who are part of crafting the description of the job, are part of evaluating, and so forth.” But there’s still a long road to go.
Does that firewall — where the person who wants to hire doesn't get to look at the process until the end — exist originally because of concerns about cronyism?
One of the interesting things about the civil service is its raison d’être — its reason for being — was supposedly a single, clear purpose: to prevent politicized hiring and patronage. That goes back to the Pendleton Civil Service Act of 1883. But it's always been a little strange that you have all of these very complex rules about every step of the process — from hiring to firing to promotion, and everything in between — to prevent political influence. We could just focus on preventing political influence, and not regulate every step of the process on the off-chance that without a clear regulation, political influence could creep in. This division [between hiring manager and applicants] is part of that general concern. There are areas where I've heard HR specialists say, "We declare that a manager is a subject matter expert, and we bring them into the process early on, we can do that." But still the division is pretty stark, and it's based on this excessive concern about patronage.
One point you flag is that the Office of Personnel Management (OPM), which is the body that thinks about personnel in the federal government, has a 300-page regulatory document for agencies on how you have to hire. There’s a remarkable amount of process.
Yes, but even that is a big change from the Federal Personnel Manual, which was the 10,000-page document that we shredded in the 1990s. In the ‘90s, OPM gave the agencies what's called “delegated examining authorities.” This says, “You, agency, have power to decide who to hire, we're not going to do the central supervision anymore. But, but, but: here's the 300-page document that dictates exactly how you have to carry out that hiring.”
So we have some decentralization, allowing managers more authority to control their own departments. But this two-level oversight — a local HR department that's ultimately being overseen by the OPM — also leads to a lot of slip ‘twixt cup and lip, in terms of how something gets implemented. If you're in the agency and you're concerned about the OPM overseeing your process, you're likely to be much more careful than you would like to be. “Yes, it's delegated to me, but ultimately, I know I have to answer to OPM about this process. I'm just going to color within the lines.”
I often cite Texas, which has no central HR office. Each agency decides how it wants to hire. In a lot of these reform states, if there is a central personnel office, it's an information clearinghouse or reservoir of models. “You can use us, the central HR office, as a resource if you want us to help you post the job, evaluate it, or help manage your processes, but you don't have to.” That's the goal we should be striving for in a lot of the federal reforms. Just make OPM a resource for the managers in the individual departments to do their thing or go independent.
Let's say I somehow get through the hiring process. You offer me a job at the Department of Transportation. What are you paying me?
This is one of the more stultified aspects of the federal civil service system. OPM has another multi-hundred-page handbook called the Handbook of Occupational Groups and Families. Inside that, you’ve got 49 different “groups and families,” like “Clerical occupations.” Inside those 49 groups are a series of jobs, sometimes dozens, like “Computer Operator.” Inside those, they have independent documents — often themselves dozens of pages long — detailing classes of positions. Then you as a manager have to evaluate these nine factors, which can each give points to each position, which decides how you get slotted into this weird Government Schedule (GS) system [the federal payscale].
Again, this is actually an improvement. Before, you used to have the Civil Service Commission, which went around staring very closely at someone over their typewriter and saying, "No, I think you should be a GS-12, not a GS-11, because someone over in the Department of Defense who does your same job is a GS-12." Now this is delegated to agencies, but again, the agencies have to listen to the OPM on how to classify and set their jobs into this 15-stage GS-classification system, each stage of which has 10 steps which determine your pay, and those steps are determined mainly by your seniority. It's a formalized step-by-step system, overwhelmingly based on just how long you've sat at your desk.
Let's be optimistic about my performance as a civil servant. Say that over my first three years, I'm just hitting it out of the park. Can you give me a raise? What can you do to keep me in my role?
Not too much. For most people, the within-step increases — those 10 steps inside each GS-level — is just set by seniority. Now there are all these quality step increases you can get, but they're very rare and they have to be documented. So you could hypothetically pay someone more, but it's going to be tough. In general, the managers just prefer to stick to seniority, because not sticking to it garners a lot of complaints. Like so much else, the goal is, "We don't want someone rewarding an official because they happen to share their political preferences." The result of that concern is basically nobody can get rewarded at all, which is very unfortunate.
We do have examples in state and federal government of what's known as broadbanding, where you have very broad pay scales, and the manager can decide where to slot someone. Say you're a computer operator, which can mean someone who knows what an Excel spreadsheet is, or someone who's programming the most advanced AI systems. As a manager in South Carolina or Florida, you have a lot of discretion to say, "I can set you 50% above the market rate of what this job technically would go for, if I think you're doing a great job."
What if you want to bring me into the Senior Executive Service (SES)? Theoretically, that sits at the top of the General Service scale. Can't you bump me up in there and pay me what you owe me?
I could hypothetically bring you in as a senior executive servant. The SES was created in the 1978 Civil Service Reform Act. The idea was, “We're going to have this elite cadre of about 8,000 individuals at the top of the federal government, whose employment will be higher-risk and higher-reward. They might be fired, and we're going to give them higher pay to compensate for that.”
Almost immediately, that did not work out. Congress was outraged at the higher pay given to the top officials and capped it. Ever since, how much the SES can get paid has been tightly controlled. As in most of the rest of the federal government, where they establish these performance pay incentives or bonuses — which do exist — they spread them like peanut butter over the whole service. To forestall complaints, everyone gets a little bit every two or three years.
That's basically what happened to the SES. Their annual pay is capped at the vice president’s salary, which is a cap for a lot of people in the federal government. For most of your GS and other executive scales, the cap is Congress's salary. [NB: This is no longer exactly true, since Congress froze its own salaries in 2009. The cap for GS (currently about $195k) is now above congressional salaries ($174k).]
One of the big problems with pay in the federal government is pay compression. Across civil service systems, the highest-skilled people tend to be paid much less than the private sector, and the lowest-skilled people tend to get paid much more. The political science reason for that is pretty simple: the median voter in America still decides what seems reasonable. To the median voter, the average salary of a janitor looks low, and the average salary of a scientist looks way too high. Hence this tendency to pay compression. Your average federal employee is probably overpaid relative to the private sector, because the lowest-skilled employees are paid up to 40% higher than the private sector equivalent. The highest-paid employees, the post-graduate skilled professionals, are paid less. That makes it hard to recruit the top performers, but it also swells the wage budget in a way that makes it difficult to talk about reform.
There's a lot of interest in this administration in making it easier to recruit talent and get rid of under-performers. There have been aggressive pushes to limit collective bargaining in the public sector. That should theoretically make it easier to recruit, but it also increases the precariousness of civil service roles. We've seen huge firings in the civil service over the last six months.
Classically, the explicit trade-off of working in the federal government was, “Your pay is going to be capped, but you have this job for life. It's impossible to get rid of you.” You trade some lifetime earnings for stability. In a world where the stability is gone, but pay is still capped, isn't the net effect to drive talent away from the civil service?
I think it's a concern now. On one level it should be ameliorated, because those who are most concerned with stability of employment do tend to be lower performers. If you have people who are leaving the federal service because all they want is stability, and they're not getting that anymore, that may not be a net loss. As someone who came out of academia and knows the wonder of effective lifetime annuities, there can be very high performers who like that stability who therefore take a lower salary. Without the ability to bump that pay up more, it's going to be an issue.
I do know that, internally, the Trump administration has made some signs they're open to reforms in the top tiers of the SES and other parts of the federal government. They would be willing to have people get paid more at that level to compensate for the increased risks since the Trump administration came in. But when you look at the reductions in force (RIFs) that have happened under Trump, they are overwhelmingly among probationary employees, the lower-level employees.
With some exceptions. If you've been promoted recently, you can get reclassified as probationary, so some high-performers got lumped in.
Absolutely. The issue has been exacerbated precisely because the RIF regulations that are in place have made the firings particularly damaging. If you had a more streamlined RIF system — which they do have in many states, where seniority is not the main determinant of who gets laid off — these RIFs could be removing the lower-performing civil servants and keeping the higher-performing ones, and giving them some amount of confidence in their tenure.
Unfortunately, the combination of large-scale removals with the existing RIF regs, which are very stringent, has demoralized some of the upper levels of the federal government. I share that concern. But I might add, it is interesting, if you look at the federal government's own figures on the total civil service workforce, they have gone down significantly since Trump came in office, but I think less than 100,000 still, in the most recent numbers that I've seen. I'm not sure how much to trust those, versus some of these other numbers where people have said 150,000, 200,000.
Whether the Trump administration or a future administration can remove large numbers of people from the civil service should be somewhat divorced from the general conversation on civil service reform. The main debate about whether or not Trump can do this centers around how much power the appropriators in Congress have to determine the total amount of spending in particular agencies on their workforce. It does not depend necessarily on, "If we're going to remove people — whether for general layoffs, or reductions in force, or because of particular performance issues — how can we go about doing that?"
My last-ditch hope to maintain a bipartisan possibility of civil service reform is to bracket, “How much power does the president have to remove or limit the workforce in general?” from “How can he go about hiring and firing, et cetera?”
I think making it easier for the president to identify and remove poor performers is a tool that any future administration would like to have.
We had this conversation sparked again with the firing of the Bureau of Labor Statistics commissioner. But that was a position Congress set up to be appointed by the President, confirmed by the Senate, and removable by the President. It’s a separate issue from civil service at large. Everyone said, “We want the president to be able to hire and fire the commissioner.” Maybe firing the commissioner was a bad decision, but that's the situation today.
Attentive listeners to Statecraft know I’m pretty critical, like you are, of the regulations that say you have to go in order of seniority. In mass layoffs, you’re required to fire a lot of the young, talented people.
But let's talk about individual firings. I've been a terrible civil servant, a nightmarish employee from day one. You want to discipline, remove, suspend, or fire me. What are your options?
Anybody who has worked in the civil service knows it's hard to fire bad performers. Whatever their political valence, whatever they feel about the civil service system, they have horror stories about a person who just couldn't be removed.
In the early 2010s, a spate of stories came out about air traffic controllers sleeping on the job. Then-transportation secretary, Ray LaHood, made a big public announcement: "I'm going to fire these three guys." After these big announcements, it turned out he was only able to remove one of them. One retired, and another had their firing reduced to a suspension.
You had another horrific story where a man was joking on the phone with friends when a plane crashed into a helicopter and killed nine people over the Hudson River. National outcry. They said, "We're going to fire this guy." In the end, after going through the process, he only got a suspension. Everyone agrees it's too hard.
The basic story is, you have two ways to fire someone. Chapter 75, the old way, is often considered the realm of misconduct: You've stolen something from the office, punched your colleague in the face during a dispute about the coffee, something illegal or just straight-out wrong. We get you under Chapter 75.
The 1978 Civil Service Reform Act added Chapter 43, which is supposed to be the performance-based system to remove someone. As with so much of that Civil Service Reform Act, the people who passed it thought this might be the beginning of an entirely different system.
In the end, lots of federal managers say there's not a huge difference between the two. Some use 75, some use 43. If you use 43, you have to document very clearly what the person did wrong. You have to put them on a performance improvement plan. If they failed a performance improvement plan after a certain amount of time, they can respond to any claims about what they did wrong. Then, they can take that process up to the Merit Systems Protection Board (MSPB) and claim that they were incorrectly fired, or that the processes weren't carried out appropriately. Then, if they want to, they can say, “Nah, I don’t like the order I got,” and take it up to federal courts and complain there. Right now, the MSPB doesn't have a full quorum, which is complicating some of the recent removal disputes.
You have this incredibly difficult process, unlike the private sector, where your boss looks at you and says, "I don't like how you're giving me the stink-eye today. Out you go." One could say that's good or bad, but, on the whole, I think the model should be closer to the private sector. We should trust managers to do their job without excessive oversight and process. That's clearly about as far from the realm of possibility as the current system, under which the estimate is 6-12 months to fire a very bad performer. The number of people who win at the Merit Systems Protection Board is still 20-30%.
This goes into another issue, which is unionization. If you're part of a collective bargaining agreement — most of the regular federal civil service is — first, you have to go with this independent, union-based arbitration and grievance procedure. You're about 50/50 to win on those if your boss tries to remove you.
So if I’m in the union, we go through that arbitration grievance system. If you win and I’m fired, I can take it to the Merit Systems Protection Board. If you win again, I can still take it to the federal courts.
You can file different sorts of claims at each part. On Chapter 43, the MSPB is supposed to be about the process, not the evidence, and you just have to show it was followed. On 75, the manager has to show by preponderance of the evidence that the employee is harming the agency. Then there are different standards for what you take to the courts, and different standards according to each collective bargaining agreement for the grievance procedure when someone is disciplined. It’s a very complicated, abstruse, and procedure-heavy process that makes it very difficult to remove people, which is why the involuntary separation rate at the federal government and most state governments is many multiples lower than the private sector.
So, you would love to get me off your team because I'm abysmal. But you have no stomach for going through this whole process and I'm going to fight it. I'm ornery and contrarian and will drag this fight out. In practice, what do managers in the federal government do with their poor performers?
I always heard about this growing up. There's the windowless office in the basement without a phone, or now an internet connection. You place someone down there, hope they get the message, and sooner or later they leave. But for plenty of people in America, that's the dream job. You just get to sit and nobody bothers you for eight hours. You punch in at 9 and punch out at 5, and that's your day. "Great. I'll collect that salary for another 10 years." But generally you just try to make life unpleasant for that person.
Public sector collective bargaining in the US is new. I tend to think of it as just how the civil service works. But until about 50 years ago, there was no collective bargaining in the public sector.
At the state level, it started with Wisconsin at the end of the 1950s. There were famous local government reforms beginning with the Little Wagner Act [signed in 1958] in New York City. Senator Robert Wagner had created the National Labor Relations Board. His son Robert F. Wagner Jr., mayor of New York, created the first US collective bargaining system at the local level in the ‘60s. In ‘62, John F. Kennedy issued an executive order which said, "We're going to deal officially with public sector unions,” but it was all informal and non-statutory.
It wasn't until Title VII of the 1978 Civil Service Reform Act that unions had a formal, statutory role in our federal service system. This is shockingly new. To some extent, that was the great loss to many civil service reformers in ‘78. They wanted to get through a lot of these other big reforms about hiring and firing, but they gave up on the unions to try to get those. Some people think that exception swallowed the rest of the rules. The union power that was garnered in ‘78 overcame the other reforms people hoped to accomplish. Soon, you had the majority of the federal workforce subject to collective bargaining.
But that's changing now too. Part of that Civil Service Reform Act said, “If your position is in a national security-related position, the president can determine it's not subject to collective bargaining.” Trump and the OPM have basically said, “Most positions in the federal government are national security-related, and therefore we're going to declare them off-limits to collective bargaining.” Some people say that sounds absurd. But 60% of the civilian civil service workforce is the Department of Defense, Veterans Affairs, and the Department of Homeland Security. I am not someone who tries to go too easy on this crowd. I think there's a heck of a lot that needs to be reformed. But it's also worth remembering that the majority of the civil service workforce are in these three agencies that Republicans tend to like a lot.
Now, whether people like Veterans Affairs is more of an open question. We have some particular laws there about opening up processes after the scandals in the 2010s about waiting lists and hospitals. You had veterans hospitals saying, "We're meeting these standards for getting veterans in the door for these waiting lists." But they were straight-up lying about those standards. Many people who were on these lists waiting for months to see a doctor died in the interim, some from causes that could have been treated had they seen a VA doctor. That led to Congress doing big reforms in the VA in 2014 and 2017, precisely because everyone realized this is a problem.
So, Trump has put out these executive orders stopping collective bargaining in all of these agencies that touch national security. Some of those, like the Environmental Protection Agency (EPA), seem like a tough sell. I guess that, if you want to dig a mine and the Chinese are trying to dig their own mine and we want the mine to go quickly without the EPA pettifogging it, maybe. But the core ones are pretty solid. So far the courts have upheld the executive order to go in place. So collective bargaining there could be reformed.
But in the rest of the government, there are these very extreme, long collective bargaining agreements between agencies and their unions. I've hit on the Transportation Security Administration(TSA) as one that's had pretty extensive bargaining with its union. When we created the TSA to supervise airport security, a lot of people said, "We need a crème de la crème to supervise airports after 9/11. We want to keep this out of union hands, because we know unions are going to make it difficult to move people around." The Obama administration said, "Nope, we're going to negotiate with the union." Now you have these huge negotiations with the unions about parking spots, hours of employment, uniforms, and everything under the sun. That makes it hard for managers in the TSA to decide when people should go where or what they should do.
One thing we've talked about on Statecraft in past episodes — for instance, with John Kamensky, who was a pivotal figure in the Clinton-Gore reforms — was this relationship between government employees and “Beltway Bandits”: the contractors who do jobs you might think of as civil service jobs.
One critique of that ‘90s Clinton-Gore push, “Reinventing Government,” was that although they shrank the size of the civil service on paper, the number of contractors employed by the federal government ballooned to fill that void. They did not meaningfully reduce the total number of people being paid by the federal government. Talk to me about the relationship between the civil service reform that you'd like to see and this army of folks who are not formally employees.
Every government service is a combination of public employees and inputs, and private employees and inputs. There's never a single thing the government does — federal, state, or local — that doesn't involve inputs from the private sector. That could be as simple as the uniforms for the janitors. Even if you have a publicly employed janitor, who buys the mop? You're not manufacturing the mops.
I understand the critique that the excessive focus on full-time employees in the 1990s led to contracting out some positions that could be done directly by the government. But I think that misses how much of the government can and should be contracted out. The basic Office of Management and Budget (OMB) statute [OMB Circular No. A-76] defining what is an essential government duty should still be the dividing line. What does the government have to do, because that is the public overseeing a process? Versus, what can the private sector just do itself?
I always cite Stephen Goldsmith, the old mayor of Indianapolis. He proposed what he called the Yellow Pages test. If you open the Yellow Pages [phone directory] and three businesses do that business, the government should not be in that business. There's three garbage haulers out there. Instead of having a formal government garbage-hauling department, just contract out the garbage.
With the internet, you should have a lot more opportunities to contract stuff out. I think that is generally good, and we should not have the federal government going about a lot of the day-to-day procedural things that don't require public input. What a lot of people didn't recognize is how much pressure that's going to put on government contracting officers at the federal level. Last time I checked there were 40,000 contracting officers. They have a lot of power. In the most recent year for which we have data, there were $750 billion in federal contracts. This is a substantial part of our economy. If you total state and local, we're talking almost 10% of our whole economy goes through government contracts. This is mind-boggling. In the public policy world, we should all be spending about 10% of our time thinking about contracting.
One of the things I think everyone recognized is that contractors should have more authority. Some of the reform that happened with people like [Steven] Kelman — who was the Office of Federal Procurement Policy head in the ‘90s under Clinton — was, "We need to give these people more authority to just take a credit card and go buy a sheaf of paper if that's what they need. And we need more authority to get contract bids out appropriately.”
The same message that animates civil service reform should animate these contracting discussions. The goal should be setting clear goals that you want — for either a civil servant or a contractor — and then giving that person the discretion to meet them. If you make the civil service more stultified, or make pay compression more extreme, you're going to have to contract more stuff out.
People talk about the General Schedule [pay scale], but we haven't talked about the Federal Wage Schedule system at all, which is the blue-collar system that encompasses about 200,000 federal employees. Pay compression means those guys get paid really well. That means some managers rightfully think, "I'd like to have full-time supervision over some role, but I would rather contract it out, because I can get it a heck of a lot cheaper."
There's a continuous relationship: If we make the civil service more stultified, we're going to push contracting out into more areas where maybe it wouldn't be appropriate. But a lot of things are always going to be appropriate to contract out. That means we need to give contracting officers and the people overseeing contracts a lot of discretion to carry out their missions, and not a lot of oversight from the Government Accountability Office or the courts about their bids, just like we shouldn't give OPM excess input into the civil service hiring process.
This is a theme I keep harping on, on Statecraft. It’s counterintuitive from a reformer's perspective, but it’s true: if you want these processes to function better, you're going to have to stop nitpicking. You're going to have to ease up on the throttle and let people make their own decisions, even when sometimes you're not going to agree with them.
This is a tension that's obviously happening in this administration. You've seen some clear interest in decentralization, and you've seen some centralization. In both the contract and the civil service sphere, the goal for the central agencies should be giving as many options as possible to the local managers, making sure they don't go extremely off the rails, but then giving those local managers and contracting officials the ability to make their own choices. The General Services Administration (GSA) under this administration is doing a lot of government-wide acquisition contracts. “We establish a contract for the whole government in the GSA. Usually you, the local manager, are not required to use that contract if you want computer services or whatever, but it's an option for you.”
OPM should take a similar role. "Here's the system we have set up. You can take that and use it as you want. It's here for you, but it doesn't have to be used, because you might have some very particular hiring decisions to make.” Just like there shouldn't be one contracting decision that decides how we buy both a sheaf of computer paper and an aircraft carrier, there shouldn't be one hiring and firing process for a janitor and a nuclear physicist. That can't be a centralized process, because the very nature of human life is that there's an infinitude of possibilities that you need to allow for, and that means some amount of decentralization.
I had an argument online recently about New York City’s “buy local” requirement for certain procurement contracts. When they want to build these big public toilets in New York City, they have to source all the toilet parts from within the state, even if they’re $200,000 cheaper in Portland, Oregon.
I think it's crazy to ask procurement and contracting to solve all your policy problems. Procurement can’t be about keeping a healthy local toilet parts industry. You just need to procure the toilet.
This is another area where you see similar overlap in some of the civil service and contracting issues. A lot of cities have residency requirements for many of their positions. If you work for the city, you have to live inside the city. In New York, that means you've got a lot of police officers living on Staten Island, or right on the line of the north side of the Bronx, where they're inches away from Westchester. That drives up costs, and limits your population of potential employees.
One of the most amazing things to me about the Biden Bipartisan Infrastructure Law was that it encouraged contracting officers to use residency requirements: “You should try to localize your hiring and contracting into certain areas.” On a national level, that cancels out. If both Wyoming and Wisconsin use residency requirements, the net effect is not more people hired from one of those states! So often, people expect the civil service and contracting to solve all of our ills and to point the way forward for the rest of the economy on discrimination, hiring, pay, et cetera. That just leads to, by definition, government being a lot more expensive than the private sector.
Over the next three and a half years, what would you like to see the administration do on civil service reform that they haven't already taken up?
I think some of the broad-scale layoffs, which seem to be slowing down, were counterproductive. I do think that their ability to achieve their ends was limited by the nature of the reduction-in-force regulations, which made them more counterproductive than they had to be. That's the situation they inherited. But that didn't mean you had to lay off a lot of people without considering the particular jobs they were doing now.
Yeah. There are also debates obviously, within the administration, between DOGE and Russ Vought [director of the OMB] and some others on this. Some things, like the Schedule Policy/Career — which is the revival of Schedule F in the first Trump administration — are largely a step in the right direction. Counter to some of the critics, it says, “You can remove someone if they're in a policymaking position, just like if they were completely at-will. But you still have to hire from the typical civil service system.” So, for those concerned about politicization, that doesn't undermine that, because they can't just pick someone from the party system to put in there. I think that's good.
They recently had a suitability requirement rule that I think moved in the right direction. That says, “If someone's not suitable for the workforce, there are other ways to remove them besides the typical procedures.” The ideal system is going to require some congressional input: it’s to have a decentralization of hiring authority to individual managers. Which means the OPM — now under Scott Kupor, who has finally been confirmed — saying, "The OPM is here to assist you, federal managers. Make sure you stay within the broad lanes of what the administration's trying to accomplish. But once we give you your general goals, we're going to trust you to do that, including hiring.”
I've mentioned it a few times, but part of the Chance to Compete Act — which was mentioned in one of Trump's Day One executive orders, people forget about this — was saying, “Implement the Chance to Compete Act to the maximum extent of the law.” Bring more subject-matter expertise into the hiring process, allow more discretion for managers and input into the hiring process. I think carrying that bipartisan reform out is going to be a big step, but it's going to take a lot more work.
DOGE could have made USAID much more accountable and efficient by listening to people like Dean, and reformers of foreign aid should think carefully about Dean’s criticisms of USAID, and his points for how to make foreign aid not just resilient but politically popular in the long term.
This is a long conversation: you can jump to a specific section with the index above. If you just want to hear about Dean’s experience with DOGE, you can click here or go to the 45-minute mark in the audio. And if you want my abbreviated summary of the conversation, see these twoTwitter threads. But I think the full conversation is enlightening, especially if you want to understand the American foreign aid system.Thanks to Harry Fletcher-Wood for his judicious edits.
Our past coverage of USAID
Dean, I'm curious about the limits of your authority. What can the Chief Economist of USAID do? What can they make people do?
There had never been an Office of the Chief Economist before. In a sense, I was running a startup, within a 13,000-employee agency that had fairly baked-in, decentralized processes for doing things.
Congress would say, "This is how much to spend on this sector and these countries." What you actually fund was decided by missions in the individual countries. It was exciting to have that purview across the world and across many areas, not just economic development, but also education, social protection, agriculture. But the reality is, we were running a consulting unit within USAID, trying to advise others on how to use evidence more effectively in order to maximize impact for every dollar spent.
We were able to make some institutional changes, focused on basically a two-pronged strategy. One, what are the institutional enablers — the rules and the processes for how things get done — that are changeable? And two, let's get our hands dirty working with the budget holders who say, "I would love to use the evidence that's out there, please help guide us to be more effective with what we're doing."
There were a lot of willing and eager people within USAID. We did not lack support to make that happen. We never would've achieved anything, had there not been an eager workforce who heard our mission and knocked on our door to say, "Please come help us do that."
What do you mean when you say USAID has decentralized processes for doing things?
Earmarks and directives come down from Congress. [Some are]about sector: $1 billion dollars to spend on primary school education to improve children's learning outcomes, for instance. The President’s Emergency Plan for AIDS Relief (PEPFAR) [See our interview with former PEPFAR lead Mark Dybul]is one of the biggest earmarks to spend money specifically on specific diseases. Then there's directives that come down about how to allocate across countries.
Those are two conversations I have very little engagement on, because some of that comes from Congress. It’s a very complicated, intertwined set of constraints that are then adhered to and allocated to the different countries. Then what ends up happening is — this is the decentralized part — you might be a Foreign Service Officer (FSO) working in a country, your focus is education, and you’re given a budget for that year from the earmark for education and told, "Go spend $80 million on a new award in education." You’re working to figure out, “How should we spend that?” There might be some technical support from headquarters, but ultimately, you're responsible for making those decisions. Part of our role was to help guide those FSOs towards programs that had more evidence of effectiveness.
Could you talk more about these earmarks? There's a popular perception that USAID decides what it wants to fund. But these big categories of humanitarian aid, or health, or governance, are all decided in Congress. Often it's specific congressmen or congresswomen who really want particular pet projects to be funded.
That's right. And the number that I heard is that something in the ballpark of 150-170% of USAID funds were earmarked. That might sound horrible, but it's not.
How is that possible?
Congress double-dips, in a sense: we have two different demands. You must spend money on these two things. If the same dollar can satisfy both, that was completely legitimate. There was no hiding of that fact. It's all public record, and it all comes from congressional acts that create these earmarks. There's nothing hidden underneath the hood.
Will you give me examples of double earmarking in practice? What kinds of goals could you satisfy with the same dollar?
There’s an earmark for Development Innovation Ventures (DIV) to do research, and an earmark for education. If DIV is going to fund an evaluation of something in the education space, there's a possibility that that can satisfy a dual earmark requirement. That's the kind of thing that would happen. One is an earmark for a process: “Do really careful, rigorous evaluations of interventions, so that we learn more about what works and what doesn't." And another is, "Here's money that has to be spent on education." That would be an example of a double dip on an earmark.
And within those categories, the job of Chief Economist was to help USAID optimize the funding? If you're spending $2 billion on education, “Let's be as effective with that money as possible.”
That's exactly right. We had two teams, Evidence Use and Evidence Generation. It was exactly what it sounds like. If there was an earmark for $1 billion dollars on education, the Evidence Use team worked to do systematic analysis: “What is the best evidence out there for what works for education for primary school learning outcomes?” Then, “How can we map that evidence to the kinds of things that USAID funds? What are the kinds of questions that need to be figured out?”
It’s not a cookie-cutter answer. A systematic review doesn’t say, "Here's the intervention. Now just roll it out everywhere." We had to work with the missions — with people who know the local area — to understand, “What is the local context? How do you appropriately adapt this program in a procurement and contextualize it to that country, so that you can hire people to use that evidence?”
Our Evidence Generation team was trying to identify knowledge gaps where the agency could lead in producing more knowledge about what works and what doesn't. If there was something innovative that USAID was funding, we were huge advocates of, "Great, let's contribute to the global public good of knowledge, so that we can learn more in the future about what to do, and so others can learn from us. So let's do good, careful evaluations."
Being able to demonstrate what good came of an intervention also serves the purpose of accountability. But I've never been a fan of doing really rigorous evaluations just for the sake of accountability. It could discourage innovation and risk-taking, because if you fail, you'd be seen as a failure, rather than as a win for learning that an idea people thought was reasonable didn't turn out to work. It also probably leads to overspending on research, rather than doing programs. If you're doing something just for accountability purposes, you're better off with audits. "Did you actually deliver the program that you said you would deliver, or not?"
Awards over $100 million dollars did go through the front office of USAID for approval. We added a process — it was actually a revamped old process — where they stopped off in my office. We were able to provide guidance on the cost-effectiveness of proposals that would then be factored into the decision on whether to proceed. When I was first trying to understand Project 2025, because we saw that as a blueprint for what changes to expect, one of the changes they proposed was actually that process. I remember thinking to myself, "We just did that. Hopefully this change that they had in mind when they wrote that was what we actually put in place." But I thought of it as a healthy process that had an impact, not just on that one award, but also in helping set an example for smaller awards of, “This is how to be more evidence-based in what you're doing.”
[Further reading: Here’s a position paper Karlan’s office at USAID put out in 2024 on how USAID should evaluate cost-effectiveness.]
You’ve also argued that USAID should take into account more research that has already been done on global development and humanitarian aid. Your ideal wouldn't be for USAID to do really rigorous research on every single thing it does. You can get a lot better just by incorporating things that other people have learned.
That's absolutely right. I can say this as a researcher: to no one’s surprise, it's more bureaucratic to work with the government as a research funder than it is to work with foundations and nimble NGOs. If I want to evaluate a particular program, and you give me a choice of who the funder should be, the only reason I would choose government is if it had a faster on-ramp to policy by being inside.
The people who are setting policy should not be putting more weight on evidence that they paid for. In fact, one of the slogans that I often used at USAID is, "Evidence doesn't care who pays for it." We shouldn't be, as an agency, putting more weight on the things that we evaluated vs. things that others evaluated without us, and that we can learn from, mimic, replicate, and scale.
We — and the we here is everyone, researchers and policymakers — put too much weight on individual studies, in a horrible way. The first to publish on something gets more accolades than the second, third and fourth. That's not healthy when it comes to policy. If we put too much weight on our own evidence, we end up putting too much weight on individual studies we happen to do. That's not healthy either.
That was one of the big pieces of culture change that we tried to push internally at USAID. We had this one slide that we used repeatedly that showed the plethora of evidence out there in the world compared to 20 years ago. A lot more studies are now usable. You can aggregate that evidence and form much better policies.
You had political support to innovate that not everybody going into government has. On the other hand, USAID is a big, bureaucratic entity. There are all kinds of cross-pressures against being super-effective per dollar spent. In doing culture change, what kinds of roadblocks did you run into internally?
We had a lot of support and political cover, in the sense that the political appointees — I was not a political appointee — were huge fans. But political appointees under Republicans have also been huge fans of what we were doing. Disagreements are more about what to do and what causes to choose. But the basic idea of being effective with your dollars to push your policy agenda is something that cuts across both sides.
In the days leading up to the inauguration, we were expecting to continue the work we were doing. Being more cost-effective was something some of the people who were coming in were huge advocates for. They did make progress under Trump I in pushing USAID in that direction. We saw ourselves as able to help further that goal. Obviously, that's not the way it played out, but there isn't really anything political about being more cost-effective.
We’ll come back to that, but I do want to talk about the 2.5 years you spent in the Biden administration. USAID is full of people with all kinds of incentives, including some folks who were fully on board and supportive. What kinds of challenges did you have in trying to change the culture to be more focused on evidence and effectiveness?
There was a fairly large contingent of people who welcomed us, were eager, understood the space that we were coming from and the things that we wanted, and greeted us with open arms. There's no way we would've accomplished what we accomplished without that. We had a bean counter within the Office of the Chief Economist of moving about $1.7 billion towards programs that were more effective or had strong evaluations. That would've been $0 had there not been some individuals who were already eager and just didn't have the path for doing it.
People can see economists as people who are going to come in negative and a bit dismal — the dismal science, so to speak. I got into economics for a positive reason. We tried as often as possible to show that with an economic lens, we can help people achieve their goals better, period. We would say repeatedly to people, "We're not here to actually make the difficult choices: to say whether health, education, or food security is the better use of money. We're here to accept your goal and help you achieve more of it for your dollar spent.” We always send a very disarming message: we're there simply to help people achieve their goals and to illuminate the trade-offs that naturally exist.
Within USAID, you have a consensus-type organization. When you have 10 people sitting around a room trying to decide how to spend money towards a common goal, if you don't crystallize the trade-offs between the various ideas being put forward, you end up seeing a consensus built: that everybody gets a piece of the pie. Our way of trying to shift the culture is to take those moments and say, "Wait a second. All 10 might be good ideas relative to doing nothing, but they can't all be good relative to each other. We all share a common goal, so let's be clear about the trade-offs between these different programs. Let's identify the ones that are actually getting you the most bang for your buck."
Can you give me an example of what those trade-offs might be in a given sector?
Sure. Let's take social protection, what we would call the Humanitarian Nexus development space. It might be working in a refugee area — not dealing with the immediate crisis, but one, two, five, or ten years later — trying to help bring the refugees into a more stable environment and into economic activities. Sometimes, you would see some cash or food provided to households. The programs would all have the common goal of helping to build a sustainable livelihood for households, so that they can be more integrated into the local economy. There might be programs providing water, financial instruments like savings vehicles, and supporting vocational education. It'd be a myriad of things, all on this focused goal of income-generating activity for the households to make them more stable in the long run.
Often, those kinds of programs doing 10 different things did not actually lead to an observable impact over five years. But a more focused approach has gone through evaluations: cash transfers. That's a good example where “reducing” doesn't always mean reduce your programs just to one thing, but there is this default option of starting with a base case: “What does a cash transfer generate?"
And to clarify for people who don't follow development economics, the cash transfer is just, “What if we gave people money?”
Sometimes it is just that. Sometimes it's thinking strategically, “Maybe we should do it as a lump sum so that it goes into investments. Maybe we should do it with a planning exercise to make those investments.” Let's just call it “cash-plus,” or “cash-with-a-little-plus,” then variations of that nature. There's a different model, maybe call it, “cash-plus-plus,” called the graduation model. That has gone through about 30 randomized trials, showing pretty striking impacts on long-run income-generating activity for households. At its core is a cash transfer, usually along with some training about income-generating activity — ideally one that is producing and exporting in some way, even a local export to the capital — and access to some form of savings. In some cases, that's an informal savings group, with a community that comes and saves together. In some cases, it's mobile money that's the core. It's a much simpler program, and it's easier to do it at scale. It has generated considerable, measured, repeatedly positive impacts, but not always. There's a lot more that needs to be learned about how to do it more effectively.
[Further reading: Here’s another position paper from Karlan’s team at USAID on benchmarking against cash transfers.]
One of your recurring refrains is, “If we're not sure that these other ideas have an impact, let's benchmark: would a cash-transfer model likely give us more bang for our buck than this panoply of other programs that we're trying to run?”
The idea of having a benchmark is a great approach in general. You should always be able to beat X. X might be different in different contexts. In a lot of cases, cash is the right benchmark.
Go back to education. What's your benchmark for improving learning outcomes for a primary school? Cash transfer is not the right benchmark. The evidence that cash transfers will single-handedly move the needle on learning outcomes is not that strong. On the other hand, a couple of different programs — one called Teaching at the Right Level, another called structured pedagogy — have proven repeatedly to generate very strong impacts at a fairly modest cost. In education, those should be the benchmark. If you want to innovate, great, innovate. But your goal is to beat those. If you can beat them consistently, you become the benchmark. That's a great process for the long run. It’s very much part of our thinking about what the future of foreign aid should look like: to be structured around that benchmark.
Let's go back to those roundtables you described, where you're trying to figure out what the intervention should be for a group of refugees in a foreign country. What were the responses when you’d say, “Look, if we're all pulling in the same direction, we have to toss out the three worst ideas”?
One of the challenges is the psychology of ethics. There’s probably a word for this, but one of the objections we would often get was about the scale of a program for an individual. Someone would argue, "But this won't work unless you do this one extra thing." That extra thing might be providing water to the household, along with a cash transfer for income-generating activity, financial support, and bank accounts. Another objection would be that, "You also have to provide consumption and food up to a certain level."
These are things that individually might be good, relative to nothing, or maybe even relative to other water approaches or cash transfers. But if you’re focused on whether to satisfy the household's food needs, or provide half of what's needed — if all you're thinking about is the trade-off between full and half — you immediately jump to this idea that, "No, we have to go full. That's what's needed to help this household." But if you go to half, you can help more people. There's an actual trade-off: 10,000 people will receive nothing because you're giving more to the people in your program.
The same is true for nutritional supplements. Should you provide 2,000 calories a day, or 1,000 calories a day to more people? It's a very difficult conversation on the psychology of ethics. There's this idea that people in a program are sacrosanct, and you must do everything you can for them. But that ignores all the people who are not being reached at all.
I would find myself in conversations where that's exactly the way I would try to put it. I would say, "Okay, wait, we have the 2,000,000 people that are eligible for this program in this context. Our program is only going to reach 250,000. That's the reality. Now, let's talk about how many people we’re willing to leave untouched and unhelped whatsoever." That was, at least to me, the right way to frame this question. Do you go very intense for fewer people or broader support for more people?
Did that help these roundtables reach consensus, or at least have a better sense of what things are trading off against each other?
I definitely saw movement for some. I wouldn't say it was uniform, and these are difficult conversations. But there was a lot of appetite for this recognition that, as big as USAID was, it was still small, relative to the problems being approached. There were a lot of people in any given crisis who were being left unhelped. The minute you’re able to help people focus more on those big numbers, as daunting as they are, I would see more openness to looking at the evidence to figure out how to do the most good with the resources we have?” We must recognize these inherent trade-offs, whether we like it or not.
Back in 2023, you talked to Dylan Matthews at Vox — it's a great interview — about how it’s hard to push people to measure cost-effectiveness, when it means adding another step to a big, complicated bureaucratic process of getting aid out the door. You said,
"There are also bandwidth issues. There's a lot of competing demands. Some of these demands relate to important issues on gender environment, fairness in the procurement process. These add steps to the process that need to be adhered to. What you end up with is a lot of overworked people. And then you're saying, ‘Here's one more thing to do.’”
Looking back, what do you think of those demands on, say, fairness in the procurement process?
Given that we're going to be facing a new environment, there probably are some steps in the process that — hopefully, when things are put back in place in some form — someone can be thinking more carefully about. It's easier to put in a cleaner process that avoids some of these hiccups when you start with a blank slate.
Having said that, it's also going to be fewer people to dole out less money. There's definitely a challenge that we're going to be facing as a country, to push out money in an effective way with many fewer people for oversight. I don't think it would be accurate to say we achieved this goal yet, but my goal was to make it so that adding cost-effectiveness was actually a negative-cost addition to the process. [We wanted] to do it in a way that successfully recognized that it wasn't a cookie-cutter solution from up top for every country. But [our goal was that] the work to contextualize in a country actually simplified the process for whoever's putting together the procurement docs and deciding what to put in them. I stand by that belief that if it's done well, we can make this a negative-cost process change.
I just want to push a little bit. Would you be supportive of a USAID procurement and contracting process that stripped out a bunch of these requirements about gender, environment, or fairness in contracting? Would that make USAID a more effective institution?
Some of those types of things did serve an important purpose for some areas and not others. The tricky thing is, how do you set up a process to decide when to do it, when not? There's definitely cases where you would see an environmental review of something that really had absolutely nothing to do with the environment. It was just a cog in the process, but you have to have a process for deciding the process. I don't know enough about the legislation that was put in place on each of these to say, “Was there a better way of deciding when to do them, when not to do them?” That is not something that I was involved in in a direct way. "Let's think about redoing how we introduce gender in our procurement process" was never put on the table.
On gender, there's a fair amount of evidence in different contexts that says the way of dealing with a gender inequity is not to just take the same old program and say, "We're now going to do this for women." You need to understand something more about the local context. If all you do is take programs and say, "Add a gender component," you end up with a lot of false attribution, and you don't end up being effective at the very thing that the person [leading the program] cares to do.
In that Voxinterview, your host says, "USAID relies heavily on a small number of well-connected contractors to deliver most aid, while other groups are often deterred from even applying by the process’s complexity." He goes on to say that the use of rigorous evaluation methods like randomized controlled trials is the exception, not the norm.
On Statecraft, we talked to Kyle Newkirk, who ran USAID procurement in Afghanistan in the late 2000s, about the small set of well-connected contractors that took most of the contracts in Afghanistan. Often, there was very little oversight from USAID, either because it was hard to get out to those locations in a war-torn environment, or because the system of accountability wasn't built there.
Did you talk to people about lessons learned from USAID operating in Afghanistan?
No. I mean, only to the following extent: The lesson learned there, as I understand it, wasn't so much about the choice on what intervention to fund, it was procurement: the local politics and engagement with the governments or lack thereof. And dealing with the challenge of doing work in a context like that, where there's more risk of fraud and issues of that nature.
Our emphasis was about the design of programs to say, “What are you actually going to try to fund?” Dealing with whether there's fraud in the execution would fall more under the Inspector General and other units. That's not an area that we engaged in when we would do evaluation.
This actually gets to a key difference between impact evaluations and accountability. It's one of the areas where we see a lot of loosey-goosey language in the media reporting and Twitter. My office focused on impact evaluation. What changed in the world because of this intervention, that wouldn’t otherwise have changed? By “change in the world,” we are making a causal statement. That's setting up things like randomized controlled trials to find out, “What was the impact of this program?” It does provide some accountability, but it really should be done to look forward, in order to know, “Does this help achieve the goals we have in mind?” If so, let's learn that, and replicate it, scale it, do it again.
If you're going to deliver books to schools, medicine to health clinics, or cash to people, and you’re concerned about fraud, then you need to audit that process and see, “Did the books get to the schools, the medicine to the people, the cash to the people?” You don't need to ask, "Did the medicine solve the disease?" There's been studies already. There's a reason that medicine was being prescribed. Once it's proven to be an effective drug, you don't run randomized trials for decades to learn what you already know. If it's the prescribed drug, you just prescribe the drug, and do accountability exercises to make sure that the drugs are getting into the right hands and there isn't theft or corruption along the way.
I think it's a very intuitive thing. There's a confusion that often takes place in social science, in economic or education interventions. They somehow forget that once we know that a certain program generates a certain positive impact, we no longer need to track continuously to find out what happens. Instead, we just need to do accountability to make sure that the program is being delivered as it was designed, tested, and shown to work.
There are all these criticisms — from the waste, fraud, and corruption perspective — of USAID working with a couple of big contractors. USAID works largely through these big development organizations like Chemonics. Would USAID dollars be more effective if it worked through a larger base of contractors?
I don't think we know. There's probably a few different operating models that can deliver the same basic intervention. We need to focus on, ”What actually are we doing on the ground? What is it that we want the recipients of the program to receive, hear, or do?” and then think backwards from there: "Who's the right implementer for this?" If there's an implementer who is much more expensive for delivering the same product, let's find someone who's more cost-effective.
It’s helpful to break cost-effective programming into two things: the intervention itself and what benefits it accrues, and the cost for delivering that. Sometimes the improvement is not about the intervention, it's about the delivery model. Maybe that’s what you're saying: “These players were too few, too large, and they had a grab on the market, so that they were able to charge too much money to deliver something that others were equally able to do at lower cost." If that's the case, that says, "We should reform our procurement process,” because the reason you would see that happen is they were really good at complying with requirements that came at USAID from Congress. You had an overworked workforce [within USAID] that had to comply with all these requirements. If you had a bid between two groups, one of which repeatedly delivered on the paperwork to get a good performance evaluation, and a new group that doesn't have that track record, who are you going to choose? That's how we ended up where we are.
My understanding of the history is that it comes from a push from Republicans in the ‘80s, from [Senator] Jesse Helms, to outsource USAID efforts to contractors. So this is not a left-leaning thing. I wouldn't say it is right-leaning either. It was just a decision made decades ago. You combine that with the bureaucratic requirements of working with USAID, and you end up with a few firms and nonprofits skilled at dealing with it.
It's definitely my impression that at various points in American history, different partisans are calling for insourcing or for outsourcing. But definitely, I think you're right that the NGO cluster around USAID does spring up out of a Republican push in the eighties.
We talked to John Kamensky recently, who was on Al Gore's predecessor to DOGE in the ‘90s.
I listened to this, yeah.
I'm glad to hear it! I’m thinking of it because they also pushed to cut the workforce in the mid-90s and outsource federal functions.
Earlier, you mentioned a slide that showed what we've learned in the field of development economics over the past 20 years. Will you narrate that slide for me?
Let me do two slides for you. The slide that I was picturing was a count of randomized controlled trials in development that shows a fairly exponential growth. The movement started in the mid-to-late 1990s, but really took off in the 2000s. Even just in the past 10 years, it's seen a considerable increase. There's about 4-5,000 randomized controlled trials evaluating various programs of the kind USAID funds.
That doesn't tell you the substance of what was learned. Here's an example of substance, which is cash transfers: probably the most studied intervention out there. We have a meta-analysis that counted 115 studies. That's where you start having a preponderance of evidence to be able to say something concrete. There's some variation: you get different results in different places; targeting and ways of doing it vary. A good systematic analysis can help tease out what we can say, not just about the effect of cash, but also how to do it and what to expect, depending on how it's done. Fifteen years ago, when we saw the first few come out, you just had, "Oh, that's interesting. But it's a couple of studies, how do you form policy around that?” With 115, we can say so much more.
What else have we learned about development that USAID operators in the year 2000 would not have been able to act upon?
Think about the development process in two steps. One is choosing good interventions; the other is implementing them well. The study of implementation is historically underdone. The challenge that we face — this is an area I was hoping USAID could make inroads on — was, studying a new intervention might be of high reward from an academic perspective. But it’s a lot less interesting to an academic to do much more granular work to say, "That was an interesting program that created these groups [of aid recipients]; now let's do some further knock-on research to find out whether those groups should be made of four, six, or ten people.” It's going to have a lower reward for the researcher, but it’s incredibly important.
It's equivalent to the color of the envelope in direct marketing. You might run tests — if this were old-style direct marketing — as to whether the envelope should be blue or red. You might find that blue works better. Great, but that's not interesting to an academic. But if you run 50 of these, on a myriad of topics about how to implement better, you end up with a collection of knowledge that is moving the needle on how to achieve more impact per dollar.
That collection is not just important for policy: it also helps us learn more about the development process and the bottlenecks for implementing good programs. As we’re seeing more digital platforms and data being used, [refining implementation]is more possible compared to 20 years ago, where most of the research was at the intervention level: does this intervention work? That's an exciting transition. It's also a path to seeing how foreign aid can help in individual contexts, [as we] work with local governments to integrate evidence into their operations and be more efficient with their own resources.
There's an argument I’ve seen a lot recently: we under-invest in governance relative to other foreign aid goals. If we care about economic growth and humanitarian outcomes, we should spend a lot more on supporting local governance. What do you make of that claim?
I agree with it actually, but there's a big difference between recognizing the problem and seeing what the tool is to address it. It's one thing to say, “Politics matters, institutions matter.” There's lots of evidence to support that, including the recent Nobel Prize. It’s another beast to say, “This particular intervention will improve institutions and governance.”
The challenge is, “What do we do about this? What is working to improve this? What is resilient to the political process?” The minute you get into those kinds of questions, it's the other end of the spectrum from a cash transfer. A cash transfer has a kind of universality: Not to say you're going to get the same impact everywhere, but it's a bit easier to think about the design of a program. You have fewer parameters to decide. When you think about efforts to improve governance, you need bespoke thinking in every single place.
As you point out, it's something of a meme to say “institutions matter” and to leave it at that, but the devil is in all of those details.
In my younger years — I feel old saying that — I used to do a lot of work on financial inclusion, and financial literacy was always my go-to example. On a household level, it's really easy to show a correlation: people who are more financially literate make better financial decisions and have more wealth, etc. It's much harder to say, “How do you move the needle on financial literacy in a way that actually helps people make better decisions, absorb shocks better, build investment better, save better?” It’s easy to show that the correlation is there. It's much harder to say this program, here, will actually move the needle. That same exact problem is much more complicated when thinking about governance and institutions.
Let's talk about USAID as it stands today. You left USAID when it became clear to you that a lot of the work you were doing was not of interest to the people now running it. How did the agency end up so disconnected from a political base of support? There's still plenty of people who support USAID and would like it to be reinstated, but it was at least vulnerable enough to be tipped over by DOGE in a matter of weeks.
How did that happen?
I don't know that I would agree with the premise. I'm not sure that public support of foreign aid actually changed, I'd be curious to see that. I think aid has always been misunderstood. There are public opinion polls that show people thought 25% of the US budget was spent on foreign aid. One said, "What, do you think it should be?" People said 10%. The right answer is about 0.6%. You could say fine, people are bad at statistics, but those numbers are pretty dauntingly off. I don't know that that's changed. I heard numbers like that years ago.
I think there was a vulnerability to an effort that doesn't create a visible impact to people's lives in America, the way that Social Security, Medicare, and roads do. Foreign aid just doesn't have that luxury. I think it's always been vulnerable. It has always had some bipartisan support, because of the understanding of the bigger picture and the soft power that's gained from it. And the recognition that we are a nation built on the idea of generosity and being good to others. That was always there, but it required Congress to step in and say, "Let's go spend this money on foreign aid." I don't think that changed. What changed was that you ended up with an administration that just did not share those values.
There's this issue in foreign aid: Congress picks its priorities, but those priorities are not a ranked list of what Congress cares about. It's the combination of different interests and pressures in Congress that generates the list of things USAID is going to fund.
You could say doing it that way is necessary to build buy-in from a bunch of different political interests for the work of foreign aid. On the other hand, maybe the emergent list from that process is not the things that are most important to fund. And clearly, that congressional buy-in wasn't enough to protect USAID from DOGE or from other political pressures.
How should people who care about foreign aid reason about building a version of USAID that's more effective and less vulnerable at the same time?
Fair question. Look, I have thoughts, but by no means do I think of myself as the most knowledgeable person to say, here's the answer in the way forward. One reality is, even if Congress did object, they didn't have a mechanism in place to actually object. They can control the power of the purse the next round, but we're probably going to be facing a constitutional crisis over the Impoundment Act,to see if the executive branch can impound money that Congress spent. We'll see how this plays out. Aside from taking that to court, all Congress could do was complain.
I would like what comes back to have two things done that will help, but they don’t make foreign aid immune. One is to be more evidence-based, because then attacks on being ineffective are less strong. But the reality is, some of the attacks on its “effectiveness,” and the examples used, had nothing to do with poorly-chosen interventions. There was a slipperiness of language, calling something that they don't like “fraud” and “waste” because they didn’t like its purpose. That is very different than saying, “We actually agreed on the purpose of something, but then you implemented it in such a bad way that there was fraud and waste.” There were really no examples given of that second part. So I don't know that being more evidence-based will actually protect it, given that that wasn't the way it was really genuinely taken down.
The second is some boundaries. There is a core set of activities that have bipartisan support. How do we structure a foreign aid that is just focused on that? We need to find a way to put the things that are more controversial — whether it's the left or right that wants it — in a separate bucket. Let the team that wins the election turn that off and on as they wish, without adulterating the core part that has bipartisan support. That's the key question: can we set up a process that partitions those, so that they don't have that vulnerability? [I wrote about this problem earlier this year.]
My counter-example is PEPFAR, which had a broad base of bipartisan support. PEPFAR consistently got long-term reauthorizations from Congress, I think precisely because of the dynamic you're talking about: It was a focused, specific intervention that folks all over the political spectrum could get behind and save lives. But in government programs, if something has a big base of support, you have an incentive to stuff your pet partisan issues in there, for the same reason that “must-pass” bills get stuffed with everybody's little thing. [In 2024, before DOGE, PEPFAR’s original Republican co-sponsor came out against a long-term reauthorization, on the grounds that the Biden administration was using the program to promote abortion. Congress reauthorized PEPFAR for only one year, and that reauthorization lapsed in 2025.]
You want to carve out the things that are truly bipartisan. But does that idea have a timer attached? What if, on a long enough timeline, everything becomes politicized?
There are economic theorems about the nature of a repeated game. You can get many different equilibria in the long run. I'd like to think there's a world in which that is the answer. But we have seen an erosion of other things, like the filibuster regarding judges. Each team makes a little move in some direction, and then you change the equilibrium. We always have that risk. The goal is, how can you establish something where that doesn't happen?
It might be that what's happened is helpful, in an unintended way, to build equilibrium in the future that keeps things focused on the bipartisan aspect. Whether it's the left or the right that wants to do something that they know the other side will object to, they hold back and say, "Maybe we shouldn't do that. Because when we do, the whole thing gets blown up."
Let's imagine you're back at USAID a couple of years from now, with a broader latitude to organize our foreign aid apparatus around impact and effectiveness. What other things might we want to do — beyond measuring programs and keeping trade-offs in mind — if we really wanted to focus on effectiveness? Would we do fewer interventions and do them at larger scale?
I think we would do fewer things simpler and bigger, but I also think we need to recognize that even at our biggest, we were tiny compared to the budget of the local government. If we can do more to use our money to help them be more effective with their money, that's the biggest win to go for. That starts looking a lot like things Mark Green was putting in place [as administrator of USAID]under Trump I, under the Journey to Self-Reliance [a reorganization of USAID to help countries address development challenges themselves].
Sometimes that's done in the context of, "Let's do that for five or ten years, and then we can stop giving aid to that country." That was the way the Millennium Challenge Corporationtalked about their country selection initially. Eventually, they stopped doing that, because they realized that that was never happening. I think that's okay. As much as we might help make some changes, even if we succeed in helping the poorest country in the world use their resources better, they're still going to be poor. We're still going to be rich. There's still maybe going to be the poorest, because if we do that in the 10 poorest countries and they all move up, maybe the 11th becomes the poorest, and then we can work there. I don't think getting off of aid is necessarily the objective.
But if that was clearly the right answer, that's a huge win if we've done that by helping to prove the institutions and governance of that country so that it is rolling out better policies, helping its people better, and collecting their own tax revenue. If we can have an eye on that, then that's a huge win for foreign aid in general.
How are we supposed to be measuring the impact of soft power? I think that's a term that's not now much in vogue in DC.
There's no one answer to how to measure soft power. It's described as the influence that we gain in the world in terms of geopolitics, everything from treaties and the United Nations to access to markets; trade policy, labor policy. The basic idea of soft power manifests itself in all those different ways.
It's a more extreme version of the challenge of measuring the impact of cash transfers. You want to measure the impact of a pill that is intended to deal with disease: you measure the disease, and you have a direct measure. You want to measure the impact of cash: you have to measure a lot of different things, because you don't know how people are going to use the cash. Soft power is even further down the spectrum: you don't know exactly how aid is helping build our partnership with a country’s people and leaders. How is that going to manifest itself in the future? That becomes that much harder to do.
Having said that, there's academic studies that document everything from attitudes about America to votes at the United Nations that follow aid, and things of that nature. But it's not like there's one core set: that's part of what makes it a challenge.
I will put my cards on the table here: I have been skeptical of the idea that USAID is a really valuable tool for American soft power, for maintaining American hegemony, etc. It seems much easier to defend USAID by simply saying that it does excellent humanitarian work, and that’s valuable. The national security argument for USAID seems harder to substantiate.
I think we agree on this. You have such a wide set of things to look at, it's not hard to imagine a bias from a researcher might lead to selection of outcomes, and of the context. It's not a well-defined enough concept to be able to say, "It worked 20% of the time, and it did not in these, and the net average…" Average over what? Even though there's good case studies that show various paths where it has mattered, there's case studies that show it doesn't.
I also get nervous about an entire system that's built around [attempts to measure soft power]. It turns foreign aid into too much of a transactional process, instead of a relationship that is built on the Golden Rule, “There's people in this country that we can actually help.” Sure, there's this hope that it'll help further our national interests. But if they’re suffering from drought and famine, and we can provide support and save some lives, or we can do longer term developments and save tomorrow's lives, we ought to do that. That is a good thing for our country to do.
Yet the conversation does often come back to this question of soft power. The problem with transactional is you get exactly what you contract on: nothing more, nothing less. There's too many unknowns here, when we're dealing with country-level interactions, and engagements between countries. It needs to be about relationships, and that means supporting even if there isn't a contract that itemizes the exact quid pro quo we are getting for something.
I want to talk about what you observed in the administration change and the DOGE-ing of USAID. I think plenty of observers looked at this in the beginning and thought, “It's high time that a lot of these institutions were cleaned up and that someone took a hard look at how we spend money there.”
There was not really any looking at any of the impact of anything. That was never in the cards. There was a 90-day review that was supposed to be done, but there were no questions asked, there was no data being collected. There was nothing whatsoever being looked at that had anything to do with, “Was this award actually accomplishing what it set out to accomplish?” There was no process in which they made those kinds of evaluations on what's actually working.
You can see this very clearly when you think about what their bean counter was at DOGE: the spending that they cut. It's like me saying, "I'm going to do something beneficial for my household by stopping all expenditures on food." But we were getting something for that. Maybe we could have bought more cheaply, switched grocery stores, made a change there that got us the same food for less money. That would be a positive change. But you can't cut all your food expenditures, call that a saving, and then not have anything to eat. That's just bad math, bad economics.
But that's exactly what they were doing. Throughout the entire government, that bean counter never once said, “benefits foregone.” It was always just “lowered spending.” Some of that probably did actually have a net loss, maybe it was $100 million spent on something that only created $10 million of benefits to Americans. That's a $90 million gain. But it was recorded as $100 million. And the point is, they never once looked at what benefits were being generated from the spending. What was being asked, within USAID, had nothing to do with what was actually being accomplished by any of the money that was being spent. It was never even asked.
How do you think about risky bets in a place like USAID? It would be nice for USAID to take lots of high-risk, high-reward bets, and to be willing to spend money that will be “wasted” in the pursuit of high-impact interventions. But that approach is hard for government programs, politically, because the misses are much more salient than the successes.
This is a very real issue. I saw this the very first time I did any sort of briefing with Congress when I was Chief Economist. The question came at me, "Why doesn't USAID show us more failures?" I remember thinking to myself, "Are you willing to promise that when they show the failure, you won't punish them for the failure — that you'll reward them for documenting and learning from the failure and not doing it again?" That's a very difficult nut to crack.
There's an important distinction to make. You can have a portfolio of evidence generation, some things work and some don't, that can collectively contribute towards knowledge and scaling of effective programs. USAID actually had something like this called Development Innovation Ventures (DIV), and was in an earmark from Congress. It was so good that they raised money from the effective altruist community to further augment their pot of money.
This was strong because a lot of it was not evaluating USAID interventions. It was just funding a portfolio of evidence generation about what works, implemented by other parties. The failures aren't as devastating, because you're showing a failure of some other party: it wasn't USAID money paying for an intervention. That was a strong model for how USAID can take on some risks and do some evidence generation that is immune to the issue you just described.
If you're going to do evaluations of USAID money, the issue is very real. My overly simplistic view is that a lot of what USAID does should not be getting a highly rigorous impact evaluation. USAID should be rolling out, simple and at scale, things that have already been shown elsewhere. Let the innovation take place pre-USAID, funded elsewhere, maybe by DIV. Let smaller and more nimble nonprofits be the innovators and the documenters of what works. Then, USAID can adopt the things that are more effective and be more immune to this issue.
So yeah, there is a world that is not first-best where USAID does the things that have strong evidence already. When it comes to actual innovation, where we do need to take risks that things won't work, let that be done in a way that may be supported by USAID, but partitioned away.
I'm looking at a chart of USAID program funding in Fiscal Year 2022: the three big buckets are humanitarian, health, and governance, all on the order of $10–12 billion. Way down at the bottom, there’s $500 million for “economic growth.” What's in that bucket that USAID funds, and should that piece of the pie chart be larger?
I do think that should be larger, but it depends on how you define it. I don't say that just because I'm an economist. It goes back to the comment earlier about things that we can do to help improve local governance, and how they're using their resources. The kinds of things that might be funded would be efforts to work with local government to improve their ability to collect taxes. Or to set up efficient regulations for the banking industry, so it can grow and provide access to credit and savings. These are things that can help move the needle on macroeconomic outcomes. With that, you have more resources. That helps health and education, you have these downstream impacts. As you pointed out, the earmark on that was tiny. It did not have quite the same heartstring tug. But the logical link is huge and strong: if you strengthen the local government's financial stability, the benefits very much accrue to the Ministry of Health, the Ministry of Education, and the Ministry of Social Protection, etc.
Fighting your way out of poverty through growth is unambiguously good. You can look at many countries around the world that have grown economically, and through that, reduced poverty. But it's one thing to say that growth will alleviate poverty. It's another to say, "Here's aid money that will trigger growth." If we knew how to do that, we would've done it long ago, in a snap.
Last question. Let's say it's a clean slate at USAID in a couple years, and you have wide latitude to do things your way. I want the Dean Karlan vision for the future of USAID.
It needs to have, at the high level, a recognition that the Golden Rule is an important principle that guides our thinking on foreign aid and that we want to do unto others as we would have them do unto us. Being generous as a people is something that we pride ourselves in, our nation represents us as people, so we shouldn't be in any way shy to use foreign aid to further that aspiration of being a generous nation.
The actual way of delivering aid, I would say, three things. Simpler. Let's focus on the evidence of what works, but recognize the boundaries of that evidence and how to contextualize it. There is a strong need to understand what it means to be simpler, and how to identify what that means in specific countries and contexts.
The second is about leveraging local government, and working more to recognize that, as big as we may be, we're still going to be tiny relative to local government. If we can do more to improve how local government is using its resources, we've won.
The third is about finding common ground. There's a lot. That's one of the reasons why I've started working on a consortium with Republicans and Democrats. The things I care about are generally non-partisan. The goal is to take the aspirations that foreign aid has — about improving health, education, economic outcomes, food security, agricultural productivity, jobs, trade, whatever the case is — and how do we use the evidence that's out there to move the needle as much as we can towards those goals? A lot of topics have common ground. How do we set up a foreign aid system that stays true to the common ground? I'd like to think it's not that hard. That's what I think would be great to see happen.
Découvrez des podcasts liées à Statecraft. Explorez des podcasts avec des thèmes, sujets, et formats similaires. Ces similarités sont calculées grâce à des données tangibles, pas d'extrapolations !