Here is a question regulators rarely used to ask: not how long pilots train for cockpit emergencies, but whether that training actually resembles one. The National Transportation Safety Board, the U.S. federal transportation safety investigation agency, raised exactly that question following a 2023 Southwest Airlines bird-strike incident, calling on airlines to replace brief simulator hours with realistic, workload-heavy smoke scenarios. It is a small distinction in phrasing and a substantial one in implication – exposing an assumption encoded across many high-consequence professions: that enough accumulated experience will reliably produce expert judgment.
The assumption is not merely cultural; it is written into the gatekeeping structures institutions use to certify readiness. General surgery accreditation specifies “Total Major Cases 850” and “at least 250 operations” before a resident progresses to PGY-3. FAA commercial pilot regulations require that applicants “must log at least 250 hours of flight time.” Most U.S. states require “four years of acceptable, progressive, and verifiable work experience” for engineering licensure. These thresholds treat volume and time as baseline signals of competence – and volume is necessary, just not sufficient. Exposure can compound into adaptable expertise or harden into something that closely resembles it: procedural fluency that holds until a genuinely unfamiliar problem arrives. Whether it goes one way or the other depends on how that exposure is structured – whether any programme has deliberately built in graded difficulty, designed case variety, and a structural obligation to interrogate one’s own reasoning. Remove any one of those, and accumulated experience becomes competence that holds right up to the moment it doesn’t.
Table of Contents
When Volume Becomes a Ceiling
Most professionals are not trained in programmes built from a blank sheet of paper; they are slotted into existing rotations, lists, or project streams. In these default environments, the cases a trainee encounters are largely determined by whoever happens to present during their term – not by any deliberate plan for breadth or progression. Difficulty fluctuates with the randomness of the caseload, and learning is often measured by how independently and quickly routine work is completed. Trainees emerge fluent within the modal problems of that environment, and this fluency can look indistinguishable from expertise until something truly unfamiliar appears.
Default exposure produces a specific kind of practitioner: highly calibrated to the modal problem, undertested at the margins. Junior medical officers who rotate through services without a curated case mix build judgment tightly around the problems those services most commonly see. Early-career engineers confined to a single structure type develop deep schemas for that niche while accumulating far fewer tested reference points when confronted with a different design class. None of this is a failure of individual effort; it is the predictable outcome of exposure determined by default. The question that follows is not how much experience trainees accumulate, but what kind – and whether anyone chose that deliberately.

Deliberate Training Design
Medical education research has already established that design features, not raw exposure, drive much of training effectiveness. A Best Evidence Medical Education (BEME) systematic review of high-fidelity simulation by Issenberg and colleagues, published in Medical Teacher in 2005, found that the most effective simulation-based learning environments shared features such as explicit feedback, tight curriculum integration, repetitive practice, and a graded range of difficulty. Simply increasing the number of scenarios did not guarantee better performance. Although those findings come from simulated settings rather than spine surgery fellowships, they establish a clear principle: how experience is structured and interrogated matters at least as much as how much of it trainees accumulate.
That principle is made concrete in the Spine Surgery Fellowship directed by Dr Timothy Steel in collaboration with St Vincent’s Private Hospital and Concord Hospital. Fellows assist in approximately 500 procedures each year, spanning minimally invasive decompression, open and percutaneous fusion, disc replacement, and vertebral reconstruction. This range is not an accident of referral patterns; it is an intentional span of complexity, ensuring trainees encounter routine work and demanding reconstructions within the same programme.
Alongside this clinical volume, fellows are required to complete two research projects to final-draft standard under Steel’s supervision – a structural requirement that prevents diagnostic and technical reasoning from accumulating unexamined alongside high case volume. There is a meaningful difference between recommending that fellows engage in critical analysis and requiring them to produce it at a publishable standard: one generates intent, the other enforces accountability.
When any of those levers is absent, the gap doesn’t surface during training – it shows up when trainees first encounter problems that fall outside their familiar band, at which point the design decisions have already been made.
Three Levers, Each with a Price of Absence
Difficulty grading positions trainees at the edge of their current capability – far enough from routine to stretch judgment, close enough to it for supervised safety. Without this structure, trainees stabilise around the complexity they handle most often. They become fast and comfortable in that band. The problem is they’re rarely asked, under supervision, to manage anything just beyond it. A surgical training programme that deliberately spans routine decompression through to complex reconstruction is exercising difficulty grading as a design choice, not an accident of whatever referrals happen to arrive.
Variety brings a different dimension. A service that attracts a diverse patient population may expose trainees to a wide range of problems, but this is variety by default – the mix can narrow abruptly with demographic shifts, referral changes, or service reconfiguration. Variety by design means the programme actively ensures trainees encounter contrasting pathologies, techniques, or project types, even when this requires rotating them across services or partnering with other centres.
The recent cockpit-smoke training recommendation from the National Transportation Safety Board makes this distinction explicit at a regulatory level. Coverage in The Washington Post summarised the Board’s call for airlines to move beyond brief discussions of cockpit smoke and towards simulator sessions that realistically reproduce the visibility loss and workload surge seen in a 2023 Southwest Airlines bird-strike incident. In a May news release, the NTSB observed that “Existing training often consists only of verbal discussion of a smoke event rather than immersive simulation involving reduced visibility or elevated workload.” When training stays at the level of verbal discussion, crews accumulate hours that diverge from the sensory and cognitive demands of an actual smoke-filled cockpit – effectively rehearsing the wrong conditions. Competence at discussing smoke, it turns out, is not the same competence as managing it.
Reflective obligation addresses a third axis. When programmes formally require trainees to analyse decisions, articulate rationales, and interrogate outcomes, they develop habits that support performance when familiar patterns break down. Strip out that obligation and reasoning goes unexamined – volume accumulates, but the understanding beneath it stays opaque. A recent perspective in the Journal of Perinatology critiquing the American Board of Pediatrics’ proposed competency-based 2-year neonatal-perinatal medicine fellowship makes exactly this argument. The proposal would increase clinically focused training while dropping the scholarly work-product requirement from the core pathway and relegating scholarship to an optional third year. Survey data from 110 of 111 neonatal-perinatal medicine programmes, 76% of which opposed the model, frame analytical and scholarly work as a structural necessity, not a finishing touch. Their warning: compress the time, remove the obligation, and you risk producing practitioners with more hours and thinner habits of critical inquiry.
Why Design Remains the Exception
If deliberate architecture is so important, its relative rarity needs an honest account. One obstacle is logistical. Unlike textbook content, caseload cannot be scheduled in advance; patients and projects arrive irregularly, and services must respond to what presents. Building a curated sequence of cases on top of that reality typically requires coordination across departments and institutions. A second obstacle is cultural. Case counts, procedures logged, and years of experience are the currencies that accreditation bodies, employers, and the public most readily recognise. A programme that can point to high volumes satisfies an entrenched expectation of rigour, even when the underlying exposure was never systematically designed.
There is also an economic barrier for senior clinicians who might otherwise invest heavily in programme architecture. In many academic medical centres, clinical productivity is tracked through relative value unit systems that, as one open-access review notes, are “utilised by the vast majority of US hospitals” and “tend to overshadow academic productivity” in faculty perceptions. Subsequent research reports that productivity targets shape how physicians allocate time, prompting proposals for separate academic RVU systems to formally recognise teaching, leadership, and scholarly work. In that environment, hours spent sequencing cases, supervising reflection, or overseeing research projects are real costs relative to billable clinical work. An institution that absorbs those costs is making a visible commitment about what training is for. One that doesn’t is also making a choice – just a quieter one, whose consequences tend to appear in someone else’s hands.
What Programmes Fill Time With
Taken together, the argument is straightforward: volume fills time, but architecture determines what that time produces. Training programmes are differentiated less by how many cases or hours they log than by which cases they prioritise, how they sequence difficulty, and what forms of analysis they require. Graded difficulty without explicit reflection can extend technical competence while leaving the underlying judgment opaque, even to the trainee. Breadth without reflection can expose people to many patterns without equipping them to recognise when a new situation doesn’t match any of them. Reflection without breadth or graded challenge risks generating sophisticated analysis anchored to an unduly narrow range of experience. Each lever shapes – and limits – what the others can achieve.
What happens inside training time is a safety variable in its own right, not just a function of how long it lasts. The same question now being asked of aviation – whether training actually reproduces the conditions it’s meant to prepare for – applies with equal force to operating theatres, engineering teams, and any domain where the distance between competence and adaptability has real consequences. Fields long accustomed to measuring readiness in hours need to pay equal attention to scenario design, workload, and realism. The true measure of a programme is not how well its graduates handle the situations the curriculum anticipated, but how they respond when confronted with a case no one thought to script – when only genuinely adaptive judgment will do.
