Overfitting

I have a disc protrusion — right foraminal, L5-S1, in case you need the coordinates, and you don’t, but I’m going to give them to you anyway because that’s the kind of person I am — and a shoulder that has been sending me passive-aggressive signals for months. I haven’t had the shoulder investigated yet, because I’ve been too busy investigating orthopedic surgeons. The shoulder is patient. I am not. We make an excellent dysfunctional team.

I built what I can only describe as a competitive intelligence report on every relevant specialist in Romania. I compared training backgrounds across three countries. I counted annual arthroscopy volumes the way normal people count calories: obsessively, and with diminishing returns on happiness. I cross-referenced surgical mentorship lineages as though orthopedics were a doctoral program and I were sitting on the admissions committee — the kind of committee member who asks the third follow-up question and makes everyone silently wish for lunch. I checked whether each surgeon operates exclusively on one joint, because a surgeon who also does knees on Tuesdays cannot, in my apparently very particular worldview, be fully trusted with a shoulder. I am aware that this standard, applied consistently, would disqualify most of the medical profession. I applied it anyway.

By day three, I had my answer. Dr. Andrei Popescu: the only orthopedic surgeon in Romania whose entire practice is dedicated to the shoulder. Six years of training in Germany, refined at the Alps Surgery Institute in Annecy under Laurent Lafosse — one of the pioneers of modern arthroscopic technique. Over four hundred shoulder arthroscopies a year. The data was unambiguous. The decision was made. The rational thing to do was stop.

I then spent eleven more days confirming what I already knew. Somewhere around day ten, I caught myself comparing the number of Latarjet procedures performed by two surgeons I had already eliminated from my shortlist. I was, at that point, optimizing a decision I had already made, using criteria that couldn’t change the outcome, on candidates who weren’t candidates. If there is an Olympic event for this, I would like to represent Romania.

In the meantime, I also built similar dossiers for a knee specialist, an ankle specialist, a spine surgeon, and an ENT surgeon for an entirely separate condition. Because when you find a hammer this satisfying, everything starts looking like a nail — including body parts that are not, by any medical definition, nails.

There’s a word for this in machine learning. It’s called overfitting. I work in machine learning. The irony is not subtle, but the universe has never been accused of subtlety, and frankly I’m not in a position to judge.


Through the lens of machine learning

In machine learning, overfitting is what happens when a model learns its training data too well — not the underlying patterns, but the specific data points, noise and all. A model that has overfit performs beautifully on data it has already seen and catastrophically on anything new. It has memorized instead of understood, which is a distinction that sounds philosophical until it costs you three weeks of your life and a noticeable increase in existential doubt.

The mechanism is well-documented. Given enough parameters and enough training time, a neural network will begin fitting not just the signal but also the noise in the data. The training loss keeps dropping — every internal metric says things are improving — while the gap between training performance and real-world performance quietly widens. The model grows more confident and more wrong at the same time, which is a combination that should sound familiar to anyone who has ever been on the internet, attended a board meeting, or listened to me explain my surgeon research methodology at dinner.

The standard remedy is regularization — techniques that deliberately constrain the model. Dropout randomly silences neurons during training, forcing the network to develop redundant representations rather than relying on any single pathway. L2 regularization penalizes large weights, keeping the model from placing excessive confidence on individual features. Early stopping halts training before the model has time to memorize. All of these work. None of them feel good to the model, which is probably why I don’t apply them to myself.

What unites all regularization techniques is that they work by making the model worse at the specific task during training. You deliberately introduce noise, randomness, constraint. You accept lower performance on known data to achieve better performance on everything else. This is counterintuitive in the same way that “stop researching and just book the appointment” is counterintuitive: technically correct, emotionally unacceptable, and the subject of ongoing internal negotiations that have not yet produced a treaty.

The counterintuitive insight — and the one I keep failing to apply to my own behavior — is that the path to better understanding sometimes runs through less optimization, not more. More data doesn’t help if you’re fitting to the wrong signal. More effort doesn’t help if the effort itself has become the objective. But try explaining that to a brain with twelve browser tabs open during lunch break and a dopamine system that treats each new piece of information as a small personal victory regardless of whether it changes anything at all. I go to bed by eleven — I’m disciplined about sleep, which is apparently the one domain where I’ve successfully implemented early stopping. But the research doesn’t need the night. It fills every gap the day offers.


Through the lens of neuroscience

The brain, it turns out, has a vested chemical interest in keeping you searching. It is not a neutral research assistant. It has a commission structure.

In 2014, Costa, Tran, Turchi, and Averbeck at the National Institute of Mental Health published a study in Behavioral Neuroscience demonstrating that dopamine directly modulates novelty-seeking during decision-making. When they blocked dopamine reuptake in monkeys using a selective dopamine transporter inhibitor (GBR-12909), the animals became significantly more drawn to novel options — even when familiar options offered objectively better rewards. Crucially, the drug didn’t change how quickly the monkeys learned which options were good. It changed how much value they assigned to the sheer novelty of unexplored choices. The learning was fine. The wanting was distorted. The monkeys were, in effect, doing exactly what I do with browser tabs, except they had the excuse of being pharmaceutically induced.

A 2022 study by Ilya Monosov’s lab at Washington University School of Medicine, published in Nature Neuroscience, went further and identified the zona incerta — a small region deep in the brain — as a key controller of novelty-seeking motivation. When zona incerta neurons were disrupted, the animals’ drive to explore novel stimuli dropped measurably. The region sits at a junction between higher-order visual areas (which encode what is new and meaningful) and circuits controlling gaze and attention. It is, architecturally, a system optimized for exactly what I do between tasks at work: scanning for the next piece of information, regardless of whether it’s needed. The zona incerta does not have a concept of “enough.” Neither, apparently, do I. We have a lot in common, which is not a compliment to either of us.

This means my eleven extra days of surgeon research were not purely a personality flaw — though they were also that. They were, at least in part, a dopaminergic reward loop: each new data point triggered a small neurochemical reward for the act of finding it, independent of whether it changed the decision. The brain does not distinguish between information that alters your conclusion and information that merely confirms it. Both feel productive. Only one is. The brain’s opinion on the matter is, “More, please,” and the brain’s opinion is difficult to override because it is also the organ responsible for overriding things.

I am a regularization researcher who cannot regularize himself. There’s a dissertation in that, but writing it would probably be another form of overfitting, and I already have enough of those in progress.


Through the lens of economics

The concept of diminishing marginal returns has been a foundational principle in economics since at least the work of David Ricardo in the early nineteenth century, though the formalization is usually credited to Johann Heinrich von Thünen and later to Alfred Marshall, because economics, like most fields, has a complicated relationship with who gets credit for obvious things. The principle is straightforward: each additional unit of input produces less additional output than the previous one. The first hour of surgeon research eliminated 90% of unsuitable candidates. The second hour refined the shortlist to three. Hours three through forty added precision that was, by any honest accounting, indistinguishable from noise but felt, in the moment, like the most important work I had ever done.

George Stigler — who won the 1982 Nobel Memorial Prize in Economics — formalized the economics of information search in his 1961 paper “The Economics of Information,” published in the Journal of Political Economy. Stigler argued that a rational searcher should stop looking when the expected cost of the next search exceeds the expected value of the information it would yield. The optimal strategy is not to find the best option. It is to find the option whose expected quality minus total search cost is highest. Searching more always finds something; the question is whether that something changes anything. This is the kind of insight that wins Nobel Prizes precisely because it is simultaneously obvious and systematically ignored by every human being alive.

The theory is elegant. My behavior violates it comprehensively. Stigler’s model assumes the searcher can estimate the value of future information. But the value of the next piece of information is precisely the thing you don’t know until you’ve found it, which creates a recursive problem that feels suspiciously like a trap designed by a universe that enjoys watching economists argue. The universe may not be malicious, but it has a documented sense of humor about optimization.

There is also Herb Simon’s satisficing — an economic model of decision-making that says rational agents should find the first option that meets their criteria and stop. Simon won the 1978 Nobel in Economics for this, which means two people received humanity’s highest recognition in economics for essentially telling people to relax, and humanity has continued not to relax. His core observation: the rational response to limited information, limited time, and limited cognitive capacity is not to search harder. It is to develop heuristics — stopping rules, good-enough thresholds — that sacrifice theoretical optimality for practical robustness. In machine learning terms, Simon was describing regularization three decades before the term existed in AI. Economists just called it being reasonable, which is, in retrospect, the more honest name — and the one that nobody has yet won a prize for practicing consistently.


Through the lens of evolutionary biology

For decades, the Irish elk (Megaloceros giganteus) served as the canonical example of evolutionary overfitting — the animal equivalent of a cautionary business case study. The story was clean: sexual selection drove antler size to extremes — spans of up to 4.2 meters — until the antlers became a fatal liability in dense post-glacial forests, and the species went extinct. It was a perfect parable about optimization gone too far. It was in textbooks. It was in TED talks. It was the biological equivalent of a moral fable, and like most moral fables, it was largely wrong.

Gould’s 1974 paper in Evolution showed that Irish elk antler size was proportional to body size, following the same allometric scaling as other deer species. They weren’t disproportionately large. They were exactly the size you’d predict for a deer that big, which is a less dramatic finding and therefore one that took considerably longer to be taken seriously. More recent research — including Moen and colleagues (2008) challenging the nutritional-cost hypothesis, and Lister and Stuart’s 2019 radiocarbon analysis in Quaternary International — points to climate change during the Younger Dryas as the primary extinction driver, not the antlers. The actual cause of death was habitat change. The antlers were innocent bystanders.

The irony here is almost too on-the-nose for an article about overfitting: we overfit the story of the Irish elk to our preferred narrative about the dangers of overfitting. We saw what we wanted to see — a dramatic morality tale about excess — and ignored the data that didn’t support the plot. The actual extinction was caused by something much less narratively satisfying: a slow change in environmental conditions that made the species’ entire ecological niche unviable. We needed a villain. The antlers looked the part. Science, eventually, disagreed.

But evolution does produce genuine cases of runaway optimization — it just prefers to be less cinematic about it. The peacock’s tail, the bowerbird’s architectural obsession, the competitive elaboration of cichlid jaw morphology in African rift lakes — these are real examples of traits driven by selection pressure past the point of ecological prudence. The general pattern holds: when a system optimizes intensely for a single criterion without environmental feedback, it becomes increasingly specialized and increasingly fragile. The generalist survives disruption. The specialist thrives until the distribution shifts, at which point it doesn’t, and somebody writes a paper about it.

I recognize the pattern in myself, which I mention not because self-awareness fixes the problem — it demonstrably does not — but because it adds an ironic layer that the article seems to require at this point. I drill deeper and deeper into measurable dimensions of surgeon quality — arthroscopy counts, training pedigrees, subspecialization breadth — while quietly ignoring unmeasurable dimensions like bedside manner, temperament under pressure, or the capacity to say “I don’t know” when the MRI is ambiguous. I’m selecting on the features I can quantify, not the ones that matter most. That’s overfitting at the species level and at the personal level, and it’s the same error: optimize for what you can count, underweight what you can’t, and hope the environment doesn’t test the dimensions you ignored.


Through the lens of psychology

Barry Schwartz, a psychologist at Swarthmore College, introduced the distinction between maximizers and satisficers in his 2004 book The Paradox of Choice. Satisficers establish criteria, search until they find something that meets those criteria, and stop — a process that sounds simple until you realize it requires the one thing maximizers lack, which is the ability to stop. Maximizers search exhaustively, comparing all options, seeking the provably best choice, and then comparing some more, because the provably best choice might not be the actually best choice, and the difference between those two things is where maximizers go to live.

Schwartz’s research — supported by subsequent studies including Dar-Nimrod, Rawn, and Lehman (2009) in Personality and Individual Differences — found that maximizers consistently obtained objectively better outcomes and were consistently less satisfied with them. In one study (Iyengar, Wells, and Schwartz, 2006), maximizers who had recently graduated secured starting salaries that were 20% higher than those of satisficers. They were also significantly less happy with their jobs. They had won by every external measure and lost by every internal one. There is a lesson in this, and it is a lesson I have successfully learned and completely failed to apply, which is itself a data point about the relationship between knowledge and behavior that psychology should probably study further.

The parallel to overfitting is structural, not metaphorical. The maximizer fits their decision model to every available data point — every review, every credential, every comparison dimension. The result is a decision optimized to the specific information landscape explored, and fragile to anything outside it. One new review, one contradicting data point, and the model destabilizes. Schwartz documented a specific emotional signature of this: anticipated regret. Maximizers suffer not only from the choices they made, but from the imagined alternatives they didn’t fully explore. The ghost of the unchosen option haunts them, which is inconvenient because there are always unchosen options and the ghosts don’t take weekends off.

I recognize this with the clinical detachment of someone who is describing their own symptoms to a medical AI during a lunch break that stopped being about lunch three tabs ago. I found an exceptional surgeon. I know he is exceptional. I have the data. I also have, in the back of my mind, a small persistent voice suggesting that maybe there’s someone better, somewhere, whom I haven’t found yet — and the voice will not be satisfied by evidence, because its function is not to find the answer. Its function is to keep searching. It is a process disguised as a purpose. It doesn’t have a stopping criterion, which is, if you think about it, the psychological equivalent of a while(true) loop — and anyone who works in programming knows how those end.


Through the lens of philosophy

Aristotle’s doctrine of the mean, articulated in the Nicomachean Ethics, proposes that every virtue occupies a position between two extremes — excess and deficiency. Courage sits between cowardice and recklessness. Generosity between miserliness and profligacy. The right action is never the maximal action. It is the appropriately calibrated one. Aristotle did not, to my knowledge, have access to a laptop, but if he had, I suspect the Nicomachean Ethics would contain a chapter about browser tabs.

This is, in a precise sense, a regularization theorem. It says that optimization without constraint produces vices, not virtues. The person who maximizes courage without limit becomes reckless. The person who maximizes generosity without limit becomes profligate. The person who maximizes information-seeking without limit becomes me on a Saturday afternoon, reading about surgical approaches to rotator cuff repair for a shoulder that might not even need surgery, while the shoulder in question aches quietly as if to remind me that it was hoping for a walk, not more data.

The Stoics — particularly Epictetus in the Discourses — drew a sharper line. They divided the world into what is “up to us” (our judgments, intentions, responses) and what is not (outcomes, others’ actions, the world). Overfitting, in Stoic terms, is the error of extending your optimization effort into domains where you have no control. I can choose a surgeon carefully. I cannot control what he finds when he examines my shoulder. My excessive research was not thoroughness. It was an attempt to convert uncertainty into certainty by sheer effort — and the Stoics would point out, with the characteristic calm of people who haven’t tried to book a medical appointment in Romania, that effort and control are not the same thing.

Voltaire’s formulation — “le mieux est l’ennemi du bien,” the best is the enemy of the good — captures the dynamic in seven words, which is six words fewer than I typically need to explain anything and approximately forty fewer than I used to explain my surgeon methodology to my friends, who have stopped asking. But I’d refine even Voltaire. The problem isn’t pursuing the best. The problem is not noticing when the pursuit has become the thing you’re optimizing for. I wasn’t minimizing the probability of choosing a bad surgeon. I was minimizing the feeling of not having been thorough enough. Those are different loss functions, and I was training on the wrong one — which is exactly what overfitting is, except in this case the training data was my own anxiety and the validation set was reality, and I never bothered to check the validation set because the training loss looked so good.

The most honest philosophical observation I can make about my own overfitting is this: I already know the surgeon. The appointment is pending. The research changed nothing about the outcome and everything about how I spent three weeks of my life. If that isn’t a Sisyphean act, it’s at least Sisyphus-adjacent — and if Sisyphus had had Wi-Fi, he would never have rolled the boulder at all. He would have spent eternity comparing boulders.


Connections

Overfitting connects to Entropy — which is the opposite problem: what happens when you stop imposing order altogether. They are the two failure modes of the same system, and I have personal experience with both, sometimes in the same week. The obsessive optimizer and the courtyard gate left unlocked are reacting to the same underlying condition — the world’s tendency toward disorder — with opposite and equally ineffective strategies. One builds an elaborate lock; the other installs fake cameras. Neither distributes keys.

Overfitting connects to Fragility because the overfit system is, by definition, fragile — optimized for known conditions and helpless against unknown ones. It connects to Enough because the core question lurking behind every overfit decision is “when have I searched enough?” — a question that the overfit mind treats as rhetorical and the satisficer treats as answered. It connects to Resonance because my approach to human relationships shows the same structural error: evaluating every dimension, comparing all candidates, optimizing for the best possible match, and ending up with a dataset instead of a connection. And it connects to Noise vs. Signal because overfitting is, at its mathematical core, the failure to distinguish between the two — the moment you start fitting to noise, you’ve lost the ability to tell what’s real, in data and in life.


What I Don’t Know

I haven’t explored how information theory formalizes the optimal stopping problem — the secretary problem and its 1/e solution, which I discuss in a later piece, suggest that even the mathematically optimal strategy for knowing when to stop only works 37% of the time, which is either humbling or liberating depending on how much you trusted mathematics to save you. I haven’t engaged with the Buddhist concept of attachment (upādāna), which might frame overfitting not as an optimization error but as a form of clinging — not to the outcome, but to the activity of seeking itself — which would make my surgeon research not a decision process but a meditation practice performed incorrectly. I haven’t examined political polarization as collective overfitting — ideological purity tests could be the political equivalent, fitting a worldview so tightly to a specific set of premises that it cannot accommodate new evidence — but I haven’t convinced myself the parallel is structural rather than rhetorical, and I’d rather leave a gap than force a connection and produce exactly the kind of false pattern-matching this article is about. And I haven’t addressed what might be the most uncomfortable question: whether the analytical mind that overfits to research is an evolutionary adaptation that was useful in information-scarce environments and is now misfiring catastrophically in an era where information is infinite, the dopamine reward for finding it never diminishes, and the browser has no natural closing time.


Where I Stand

Here’s the part I almost didn’t write, since it’s less funny than the rest: I booked the appointment with Dr. Popescu on day three — the same day I found him. The eleven extra days of research happened in parallel, not instead of. I wasn’t paralyzed. I was curious. I’m someone who likes to understand how things work — surgical techniques, training lineages, the taxonomy of arthroscopic approaches — not because I need all the information to act, but because understanding is genuinely one of the things that makes me feel alive.

The overfitting spiral I described in this article is real. The temptation is real. But the version of me writing this is also the version who has learned, slowly, and through enough evidence to satisfy even a maximizer, that action — imperfect, underprepared, based on incomplete data — is the thing that actually moves life forward. Not the research. Not the analysis. Not the twelfth browser tab. The appointment.

I still overfit. I probably always will. But I’ve started treating it less as a flaw to eliminate and more as a feature to manage — the way you’d manage any powerful system that doesn’t come with a built-in off switch. The curiosity stays. The research continues. But so does the living. And the living, it turns out, does not wait for the research to finish, which is the most useful thing I’ve learned and the hardest thing I’ve practiced.

If you see something I missed, I’d like to hear about it. Though I should warn you: if you send me a new perspective, I will almost certainly research it more thoroughly than necessary, cite three sources you didn’t ask for, and then write about the experience. But I’ll also do the thing. Probably before the twelfth tab.


Written: April 2026 Version: 1.0 This is how I understand this concept today. It will change.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *