top of page

The 4pm Problem

11 minutes ago
16 min read

There is a module in my course roadmap called The psychology of decision fatigue, and the line underneath it reads "why your worst decisions happen at 4pm".


I wrote that line. I liked that line. It has the ring of something everybody already half knows, which is the most dangerous quality a sentence can have.


A desk lamp lighting an alarm clock and scattered papers late in the evening, city dark through the window.
The decision arrives at 4pm. Your judgement left at 2.

Then I went and read the evidence properly, and now I have to rewrite the module.


This is what I found, in the order I found it, because the order turns out to be the interesting part. It is a story about a chart that ev

erybody has seen, a theory that fell over, and an effect that survived both.



What you will get out of this


  • Why the most famous decision fatigue study is probably measuring something other than decision fatigue

  • What happened when 2,141 people across twenty-three laboratories tried to replicate the theory underneath it

  • The evidence that did survive, from real clinical records rather than laboratory handgrips

  • A better model for why four o'clock feels the way it does, and why it changes what you should do

  • Three questions that work better than the five techniques I was going to teach

  • A free one-page decision triage card



The study everybody cites


In 2011, three researchers published an analysis of Israeli parole hearings.


The proportion of favourable rulings started each session at around sixty-five per cent and fell steadily towards zero, then reset sharply after a food break. Plotted, it looks less like data and more like a fable. It is one of the most cited findings in behavioural science, it has appeared in nearly every business book of the last decade, and you have almost certainly heard it summarised as hungry judges deny parole.


I have used it. Everyone has used it. It is an extraordinarily good slide.


Within months of publication, two other researchers published a response pointing out something fairly fundamental about how those sessions were organised. The ordering of cases was not random. Prisoners without legal representation were scheduled together, and prisoners without legal representation receive fewer favourable rulings.


Once you know that, the elegant declining curve has a considerably duller explanation sitting right next to it.


Then a third researcher ran simulations and showed that the size of the effect was implausibly large for a psychological mechanism. To produce a swing of that magnitude, decision fatigue would have to be doing something more dramatic than almost anything else in the entire literature manages. When an effect is enormous in a field where effects are small, the most likely explanation is not that you have found something enormous. It is that you are measuring something else.


The original study is not fraudulent and the authors did nothing improper. It is a real analysis of real data. It is simply not good evidence for the thing it gets cited for, which is a distinction that took roughly a decade and several thousand citations to become widely known.



The theory underneath it


Decision fatigue rests on ego depletion, proposed in 1998. Self-control draws on a limited shared resource, so spending it on one task leaves less available for the next.


It is intuitive. It is tidy. It matches the subjective experience so neatly that it feels less like a theory and more like a description. It generated hundreds of studies.


Then the replications arrived.


Study

Scale

Result

Original ego depletion paper (1998)

Laboratory

Established the limited-resource model

Multi-lab preregistered replication (2016)

Twenty-three laboratories, 2,141 participants

Effect indistinguishable from zero

Multisite preregistered test (2021)

Multiple sites, designed with the original researchers

Essentially nothing


Chart of ego depletion effect sizes: 0.62 in a 2010 meta-analysis, 0.04 in a 2016 replication across 23 laboratories, 0.06 in a 2021 test across 36 sites, both later confidence intervals crossing zero.
It took a decade and several thousand citations for this to become widely known.

That second one is the sort of thing that stops you mid-paragraph. Twenty-three laboratories. Preregistered, so nobody could go fishing afterwards. Two thousand one hundred and forty-one people. Nothing.


And the third included the original researchers in its design, which is about as fair a hearing as a theory is ever going to get.


This is where I put the pen down and reconsidered the module, because the mechanism I had planned to teach does not appear to exist. If your willpower is a battery, nobody has managed to find the battery.


I want to be careful about how hard to push that. Failed replications do not prove an effect is absent, and the debate is live rather than closed. But you cannot in good conscience build a paid module on a construct that two large preregistered efforts could not detect, and then charge people for it.



And yet


Here is what kept the whole thing alive, and it is the interesting part.


Researchers looked at antibiotic prescribing across primary care and found that clinicians prescribed antibiotics more often for acute respiratory infections later in a clinic session, in exactly the cases where the guidelines said they should not. A separate team found the same shape in opioid prescribing, with later appointments associated with more prescriptions.


These are not undergraduates squeezing handgrips in a psychology building for course credit. These are real clinical decisions with real consequences, recorded in administrative data, across very large samples, made by highly trained people who know the guidelines perfectly well and have every professional and legal reason to follow them.


The time of day moves the decision.


So the phenomenon outlived the theory that was supposed to explain it. Something is genuinely happening in the afternoon. It is simply not a fuel tank running dry, which means most of the advice built on the fuel tank metaphor is aimed at a problem that does not exist.



A better explanation


The most convincing account I read is an opportunity cost model. Effort does not drain a reservoir. Effort feels progressively more unpleasant as the number of other things you could be attending to accumulates.


That reframes four o'clock completely.


You are not out of decision fuel. You are five hours into a queue of unfinished obligations, and every single one of them is quietly raising the price of whatever is now in front of you. What you are experiencing as tiredness is closer to an accumulating awareness of everything else.


Which explains something the battery model never could, namely why a genuinely interesting problem at half past four does not feel tiring at all, while a trivial diary clash at the same hour feels like being asked to move a wardrobe.


Two further findings sharpen it.


What you believe about willpower changes whether you show depletion effects at all. People who think of self-control as unlimited do not show the pattern in the same way. Which means "I am useless after three" may be operating as an instruction rather than an observation, and saying it out loud to your team every afternoon is not the neutral act it feels like.


And the 2008 work suggests it is making choices specifically, rather than exerting effort generally, that impairs what comes next. That distinction matters enormously if your role consists largely of choosing on other people's behalf, because it means the volume of small decisions is the load. Not the difficulty of any one of them. The number.


Forty small decisions before lunch is not a light morning. It is the heaviest possible morning, and it does not look like anything on a calendar.



The part nobody has studied


I went looking for research on the specific position of making dozens of decisions a day for somebody else, while carrying the consequences of getting it wrong without holding the authority that would normally accompany the call.


I could not find it.


The literature studies people deciding for themselves, or clinicians deciding for patients inside a professional frame that comes with professional protection. The support role, where accountability points sideways rather than downwards, does not appear to have been examined at all.


Anyone who has held that role knows it has its own particular texture. The problem is not the number of decisions, exactly. It is that each one carries a small charge of anticipated answerability to another person, and by mid-afternoon you are not tired of deciding. You are tired of being answerable.


I cannot cite that and I am not going to pretend I can. What I can say is that the opportunity cost model fits it rather well, since anticipated answerability is precisely the kind of background obligation that would raise the felt price of the next decision. And that if you recognise the feeling, you are not imagining it, you are not weak, and the reason nobody has written about it is not that it is unimportant.



Three questions instead of five techniques


I had a list of tactics ready for this module. Most of them assumed a resource that does not appear to exist, so they have gone in the bin. What is left is smaller and, I think, considerably more honest.


Three questions to ask before a late-afternoon decision: what am I tired of, does this deserve a decision at all, and how expensive is it to change my mind.
The rest of the tactics assumed a resource that does not appear to exist, so they went in the bin.

One. What am I actually tired of?


If the answer is the queue rather than the deciding, the intervention is not rest. It is closing loops.


Three small things finished before a significant decision will do more than twenty minutes with your eyes shut, because you are lowering the accumulated cost rather than topping up a tank that does not exist. Order matters more than stamina.


The practical version is unglamorous. Put the decisions with real consequence early in the day, not because you are stronger at nine, but because at nine there is less unfinished business generating background cost. And if a significant decision has to happen at four regardless, spend the ten minutes beforehand clearing three trivial open items rather than preparing harder.


That advice sounds like procrastination and is the opposite of it.


Two. Does this deserve a decision at all?


Herbert Simon gave us satisficing seventy years ago. Take the first option that clears the bar rather than searching for the best one.


Most of what crosses your desk deserves exactly that treatment, and the whole skill is sorting quickly into which is which.


Iyengar, Wells and Schwartz found that people who maximise achieve objectively better outcomes and feel considerably worse about them. That is a trade almost nobody would accept if it were offered out loud, and most of us accept it several times a day without noticing we have been asked.


A rough sorting rule that has held up: if the difference between the best available option and a good-enough one is smaller than the cost of finding the best one, stop looking. You have already won and you are now paying for the privilege of continuing.


Three. How expensive is it to change my mind later?


This is the one I would keep if I could keep only one.


At four o'clock the useful question is not what the right answer is. It is how much a reversal would cost.


If reversing is cheap, decide now, move on, and stop treating a low-stakes decision as though it deserved your best thinking. If reversing is expensive, saying you will come back to it in the morning is a completely acceptable sentence.


People in support roles say it far too rarely, usually because responsiveness has quietly become part of how they are valued. A reputation for fast answers is a very poor trade against a reputation for good ones, and nobody has ever been let go for taking a night over something that mattered.


Grid crossing cheap and expensive to reverse against 10am and 4pm, showing what to do in each case.
Bottom right is the box people in support roles refuse to use.



Download: The Decision Triage Card


One page, designed to live somewhere you will actually see it at about ten past four.


What is inside

Why it is there

The three questions, with space to work them

Because five techniques built on a resource nobody can find are worse than three questions that hold up

A loop-closing prompt for the ten minutes before a big decision

Because you are managing a queue, not a battery

The satisficing sorting rule

Because most of what crosses your desk deserves the first option that clears the bar

A reversibility grid for the decisions on your desk right now

Because how expensive it is to change your mind is the only question that matters at four o'clock


Free, no email wall. Print it, stick it somewhere visible, and use the grid weekly for a month.




What I would change about the day itself


Two structural things, if you have any influence over the shape of a diary. And if you are reading this, you almost certainly have more influence over the shape of a diary than anyone else in your organisation.


Protect one block before lunch for deciding rather than delivering, and defend it the way you would defend an external meeting. Most operational roles have precisely the reverse arrangement, where thinking gets whatever is left over after everybody else's requests have been served, which is generally the worst hour of the day and eleven minutes of it.


Separate deciding from processing. A calendar made entirely of thirty-minute slots quietly teaches you that a resourcing decision and a diary clash are the same size of event. They are not. And the accumulating cost of the small ones is exactly what makes the large one feel so expensive when it finally arrives.



Frequently asked questions


Is decision fatigue real?

Something real happens, and the time-of-day effects in clinical prescribing data are hard to explain away. The popular mechanism, a limited pool of willpower, did not survive two large preregistered replications. Both of those statements are true and they are not in conflict.

I would. If you want the point, use the prescribing studies instead. They are less dramatic, considerably more solid, and nobody in the room will have heard them.

The food-break element of the parole study is the part most undermined by the scheduling critique, so I would not build anything on it. Eat because you are hungry, which is a perfectly good reason on its own.

Then the reversibility question is the one to lean on, and closing small loops beforehand matters more rather than less. You are managing the queue, not your stamina.

Almost certainly, and nobody has studied it. The opportunity cost model fits the experience well, because anticipated answerability is exactly the kind of background obligation that raises the price of the next decision.

There is no threshold in the evidence, and anybody quoting you a number is inventing it. What the research supports is that the volume of choices matters more than the difficulty of any single one.

No. It means the specific limited-resource model of self-control did not replicate. People clearly do exert self-control. The question is what it costs and why, and the opportunity cost account currently explains the evidence better.



Where to start


Judgement and decision-making is the first of five dimensions in the Connected Leader Assessment. It is free, it takes about twelve minutes, and it will tell you whether your decision-making is the thing carrying you or the thing quietly costing you.


If you want the wider territory, including risk and judgement under pressure, that is what our Skill Sprints are for. One topic, built properly, an hour or two.


The module keeps its title. The mechanism has changed, the tactics have shrunk from five to three, and it is considerably better for having been wrong first. Which is, I suspect, the actual lesson.


Meg ✌️


What the afternoon actually does to a support role


The clinical studies measure prescribing because prescribing is recorded. Nobody records the decisions you make, which is why none of this literature is about you.


So here is the shape of it, drawn from the model rather than from data, and labelled as such.


By four o'clock you are not choosing between good and bad options. You are choosing between two acceptable options while carrying eleven unfinished obligations, three of which belong to other people and one of which will become urgent tomorrow morning whether or not you touch it today.


The decision in front of you is not harder than the one you made at ten. The queue behind it is longer, and the queue is what sets the price.


That is why the standard advice fails. Take a break, drink some water, eat something. All of it treats the symptom as depletion. None of it reduces the queue.



A worked example


Twenty past four. Somebody appears with a question that has three plausible answers and a deadline of now.


The old approach. Push through, because that is what responsiveness looks like, and because saying "not now" feels like failing at the one thing you are known for. Pick an answer. Discover on Thursday that it was the wrong one, and spend forty minutes unpicking it.


The three questions instead.


First, what am I actually tired of? If it is the queue, the ten minutes before this decision are better spent closing three small things than preparing harder for this one.


Second, does this deserve a decision at all? If any of the three answers clears the bar, take the first one and stop. You have already won and you are now paying for the privilege of continuing to look.


Third, how expensive is it to change my mind? If reversal is cheap, decide now and move on with a completely clear conscience. If reversal is expensive, park it.



What to say when you park a decision


The reason people do not park decisions is not that they lack judgement. It is that they have no sentence ready, and in the absence of a sentence the default is to answer.


Here are four that work, none of which sound like stalling.


"I want to give you a proper answer rather than a fast one. I will come back to you before ten tomorrow."

"Two of these three options are fine and one is expensive to undo. Let me check the one thing that separates them and confirm this afternoon."

"I can give you an answer now if you need one, and it will be the safe one rather than the best one. Which would you rather have?"

"Yes, and I would rather not decide that at this hour. Can it hold until the morning?"

That third one is my favourite, because it hands the choice back with the cost made explicit, and most people will take the morning once they can see the trade.



The three-day experiment


If you want to know whether any of this applies to you specifically, rather than to people in general, run this. It takes about ninety seconds a day.


  1. For three days, note the time of every decision you regret, however small. A reply sent too fast. An agreement you did not mean. An option you took because it ended the conversation.

  2. Next to each one, write what you were tired of. The deciding, or the queue.

  3. On the third evening, look at the times.


Most people find a cluster. Some find it at four, some at eleven in the morning after a long meeting, some right after lunch. The cluster is more useful than anything in this article, because it is yours.


If the cluster is real, move the decisions that matter away from it. If there is no cluster, then congratulations, the afternoon is not your problem and you can stop blaming it.


Protecting the block, when the diary is not yours


Everything above assumes you can move a decision earlier in the day. For most people reading this, the diary being discussed belongs to somebody else, and the first thing to go when it gets tight is your own thinking time, because it is the only slot with no other person attached to it.


So here is the honest version of that problem.


A block in your calendar labelled "focus" will be booked over by lunchtime. A block labelled "prep for Thursday board" will not, because it has a name attached to something somebody senior cares about. This is not a hack, it is how calendars are read.


Three things that hold up in practice.


Name the block after the work, never after the state. "Thinking time" reads as availability. "Board pack review" reads as a commitment.


Put it before the meeting it serves, not after. Preparation defended in advance is preparation. Preparation scheduled afterwards is admin, and admin moves.


Make one recurring slot untouchable and let everything else flex. One genuinely protected hour a week beats five that get eaten, and the credibility of the one is what protects it.



What this research does not say


Three things I want to be clear about, because the temptation with a piece like this is to overclaim in the opposite direction.


It does not say that self-control is a myth. People clearly exert self-control, and they clearly find it effortful. What failed to replicate was one specific model of how that works, namely a shared limited pool that runs down.


It does not say that you should ignore how you feel at four o'clock. The feeling is real and it is telling you something. It is simply telling you about your queue rather than about your reserves.


And it does not say that the parole judges study was dishonest. It says that a beautiful chart got adopted faster than it got checked, which is a story about how research travels rather than a story about the researchers.


That last one is worth holding on to, because the same thing is currently happening to at least a dozen findings you and I both repeat without checking.


What I would test next


Three things I would genuinely like somebody to study, since nobody has.


Whether the queue effect is bigger when the unfinished items belong to other people. My guess is yes, and it is only a guess.


Whether anticipated answerability behaves like an open loop. If it does, the practical implication is that closing the loop socially, by telling somebody where a thing has got to, is as useful as finishing it.


And whether the four o'clock cluster moves when somebody changes role rather than changes habits. If it does, it was never about the person.


If you run any of those, I would like to read it.



Free tools


  • The decision triage tool, the interactive tool in this post, free

  • The Decision Triage Card, the download in this post, free

  • The Connected Leader Assessment, five capability dimensions, twelve minutes, free






P.S. Want to Go Deeper on Decision-Making?


If this resonated and you want to go further, here is the evidence the piece is built on. These are the papers that shaped my thinking on decision-making, effort and what survived the replication crisis, and they are well worth your time.


Every DOI below has been checked. Where a popular claim did not survive that check, I have said so in the piece rather than quietly leaving it out.



📃 Danziger, Levav and Avnaim-Pesso (2011) The parole judges study, included because you should read the thing everybody cites. https://doi.org/10.1073/pnas.1018033108


📃 Weinshall-Margel and Shapard (2011) The response showing case ordering in those sessions was not random. https://doi.org/10.1073/pnas.1110910108


📃 Glöckner (2016) Simulations showing the magnitude of the hungry judge effect is implausible for the mechanism proposed. https://doi.org/10.1017/S1930297500004812


📃 Baumeister, Bratslavsky, Muraven and Tice (1998) The original ego depletion paper and the source of the limited-resource model. https://doi.org/10.1037/0022-3514.74.5.1252


📃 Hagger and colleagues (2016) Multi-lab preregistered replication across twenty-three laboratories and 2,141 participants, effect near zero. https://doi.org/10.1177/1745691616652873


📃 Vohs, Schmeichel and colleagues (2021) Multisite preregistered paradigmatic test of ego depletion, again finding essentially nothing. https://doi.org/10.1177/0956797621989733


📃 Vohs, Baumeister, Schmeichel and colleagues (2008) Making choices specifically, rather than exerting effort generally, impairs subsequent self-control. https://doi.org/10.1037/0022-3514.94.5.883


📃 Kurzban, Duckworth, Kable and Myers (2013) The opportunity cost model of subjective effort and task performance. https://doi.org/10.1017/S0140525X12003196


📃 Job, Dweck and Walton (2010) Beliefs about willpower moderate whether depletion effects appear at all. https://doi.org/10.1177/0956797610384745


📃 Linder, Doctor, Friedberg and colleagues (2014) Antibiotic prescribing rose later in clinic sessions, in real records at scale. https://doi.org/10.1001/jamainternmed.2014.5225


📃 Neprash and Barnett (2019) The same time-of-day pattern in opioid prescribing across primary care appointments. https://doi.org/10.1001/jamanetworkopen.2019.10373


📃 Simon (1955) The behavioural model of rational choice, and the origin of satisficing. https://doi.org/10.2307/1884852


📃 Iyengar, Wells and Schwartz (2006) Maximisers achieve better objective outcomes and feel worse about them. https://doi.org/10.1111/j.1467-9280.2006.01677.x



P.P.S. Why Do I Even Have the Nerve to Write This?


Fair question.


I am someone who has spent years working alongside remarkable executives, navigating chaos, translating vision into structure and figuring things out, sometimes beautifully and sometimes the very hard way. I have had the privilege of being the steady right hand to leaders who move fast, take risks and expect a lot. That environment has shaped me more than any classroom ever could.


I also happen to have an MBA, and I am studying psychology because people fascinate me. How we work. Why we disconnect. What actually holds us together when the pressure rises.


But really, none of that is the point.


I am here because I have lived the erosion and rebuilt from it. I have seen the cost of carrying too much and the relief that arrives when you finally stop disappearing inside your own competence.


So I share what I have learned in case it helps someone else. Take what is useful. Leave the rest. And remember this one truth that professionals like you often forget:


You are probably doing far better than you give yourself credit for.


Comments


Commenting on this post isn't available anymore. Contact the site owner for more info.
Post: Blog2_Post
bottom of page