Scale Without Surprise — Introduction

Introduction

The manager’s weekly Kobayashi Maru, competing organizational obligations, and why relationships are part of the information system.

One of the first managers I worked for who understood this job offered it to me with a warning:

One of the worst jobs in IT, it’s the equivalent of passing the Kobayashi Maru test every week.

If you are trying to recruit someone, there are probably better ways to describe the opportunity. But I appreciate that he understood what he was asking me to do, because a lot of the advice you will read about capacity management assumes that the problem can be solved once everyone supplies the right inputs. Get the requirements, measure the workload, forecast demand, acquire the capacity. There you go. Why is everyone having such a difficult time?

For readers who did not grow up with Star Trek, the Kobayashi Maru is a command exercise designed to be impossible to pass under its original rules. A ship is in distress, and going to its assistance exposes the rescuing crew to overwhelming opposition. Refusing to help means abandoning the people who need rescuing. The exercise is intended to reveal how a commander responds when there is no satisfactory outcome. James T. Kirk does not solve the exercise as given: he reprograms the simulation so that a rescue becomes possible. He changes the conditions of the test.1

That is a useful way to think about this job. You still cannot reprogram a supply shortage. The business wants infinite scale, engineers want to innovate without limits, the SRE team wants that same scale with zero downtime, and Finance wants it all to cost less. Sounds easy, right?

Everyone has an unmovable constraint. Unfortunately, a plan that satisfies all of them cannot work.

My manager understood that the work was going to involve changing some of those conditions: finding out what a customer actually needed, discovering that an apparently fixed deadline had some flexibility, or showing that buying capacity now is cheaper than running out of it. You can do the calculations correctly and still be a long way from an answer. Several of the factors can be settled only once an architecture has been finalized, or you may need to commit to a limited amount of capacity before the teams are ready to make a commitment of their own.

It is often the capacity manager who is responsible for walking teams through a series of compromises and choices, and with the arrival of AI infrastructure these choices have become both more urgent and more expensive.

This is the phrase I keep coming back to:

Cloud Capacity and Cost is less something to be managed as it is a challenge that will manage you if you approach it from the wrong angle.

By the wrong angle, I mean shifting capacity decisions earlier in the development cycle and expecting the estimates made while a system is still being built to remain valid weeks or months later, when that system moves to production. That can be a perfectly reasonable approach to one component in isolation, but, at enterprise scale, while you are delivering that isolated component, another business is changing its plans, a customer is preparing a migration you haven’t heard about, and an experiment has become an executive priority overnight.

Add Generative AI to this area, and you face the additional challenge of accelerated demand: a new requirement can be coded in days, and new operational requirements can emerge quickly. This job is understanding the choices and the tradeoffs, but it is also understanding how to manage a Kobayashi Maru.

The demonstration was last week

The book returns to one constructed case throughout. The organization and its numbers are invented, but the pattern will be familiar to anyone who has done this work. An engineering team builds a new feature on top of an existing service and demonstrates it. The business sees an opportunity, and within a week the question is no longer whether the experiment works but how soon every customer can have it. Capacity planning enters only after the enthusiasm has become an expectation.

The service in this case has enough capacity for its current customers, and enough for the rollout too, provided nothing fails. What it lacks is the capacity to take on the new work while keeping a promise already made to existing customers: that the service will stay up if part of it goes down. Additional capacity has been ordered, but it will not arrive until after the launch date.

Each group at the table sees a different problem. Finance asks why the existing resources can’t work harder. Engineering asks whether the platform can scale, and the operations team asks what happens when part of it is unavailable. The business, reasonably, wants to know how a promising demonstration turned into an infrastructure discussion. Every one of those questions is legitimate, and a meeting can spend its entire hour agreeing that each one matters and still end with exactly the shortage it started with.

Later chapters give the case its full technical treatment: measuring the service, examining demand, and calculating the gap. The more instructive part comes afterward, when it turns out that some batch work could move and the rollout could start with a smaller group of customers. Those options change the arithmetic, and they come from people who understand the customers and the cost of delay. No utilization report would have surfaced them.

It would be better if requirements arrived before anyone promised a launch, and better still if they were accurate and stable. But a capacity team that answers an urgent request with a lecture about planning discipline isn’t helping anyone. The experiment that was low-risk last week still needs support now that someone wants it in front of every customer within a month. How the request should have arrived is a conversation for later.

Helping does not mean agreeing to an impossible promise. It means understanding the problem well enough to say which conditions would have to change, what each change would cost, and who has the authority to agree to it. Kirk’s question applies: are we sure we understand the rules we are trying to satisfy? (and can we change them?)

Relationships are part of the information system

The central argument of this book is that managing capacity and cost requires you to understand how the business creates demand, and that building relationships is how much of that understanding becomes accessible. I don’t mean networking so that people will attend your meeting. I mean learning enough about their work that you can recognize the difference between an idle system, a system waiting for a customer, and a system about to become a very expensive surprise.

Telemetry will tell you that a workload has been quiet. Someone on a customer team may know that several operations are about to move onto it. The billing export will show the old environment still running, but the person managing the migration can explain why it cannot be retired next month. If you only speak to engineering, you may miss the people who already have a sophisticated view of demand, expressed in units that have nothing to do with processors or requests per second.

I am an engineer, so this isn’t an argument that engineers are incapable of understanding the business. It is an argument that the information does not automatically find its way into an engineering planning system. As an engineering function, cloud capacity and cost management often has to build relationships with the business and with Finance, and it has to be aware of how sensitive their forecasts and projections can be. A revenue projection or a customer pipeline may reflect plans that haven’t been announced or numbers that haven’t been approved, and the people who hold them are right to be careful about where they go. Nor does another team necessarily want to hand it over. If the last tentative estimate they shared became a commitment in somebody else’s Powerpoint presentation, why would they share the next one? If admitting uncertainty earned them more reporting work and no assistance, another spreadsheet request is unlikely to improve the relationship.

You have to become useful to engineering, Finance, reliability, and the business, and you have to find a way to help them understand each other. In many ways, a cloud capacity and cost team acts as a bridge. Understand what they know, what they are still trying to find out, and what you are allowed to do with the information. A customer plan can be important without being final, and you need a way to use it without making the person who told you responsible for a guarantee they never offered. That is harder than adding a column labeled “confidence” to a spreadsheet. The confidence that matters has to be built with partners across the organization, over enough conversations that they come to trust what you will do with their numbers. It also gives you a better chance of hearing about the next change before it becomes a production incident.

For all this talk of trust and relationships, the engineering and the mathematical models behind a forecast still matter, and this book spends a good deal of time on both. But a model can only work with the information it is given, and the information that changes a forecast most often arrives through relationships most engineering teams have never had a reason to build. That is the part of the capacity function that is hardest to build and easiest to neglect.

What changes at scale

I have had a version of the same conversation with people at several large companies: you mention a well-known book about how they work, and the response is, in effect, “Oh, you read that book. Just know that it doesn’t represent the whole company.” I am paraphrasing recurring conversations here with employees of large tech companies, not quoting one person. The useful part is what follows, when they begin explaining the real variations in approach you could not have inferred from the book.

The conversation almost always begins with, “Oh, right, you read that book, well… it’s not that simple.”

What you learn is that a large company rarely manages capacity in just one way. A central strategy helps, and I am not proposing that every team invent its own discipline, but a company that supports different businesses, markets, technologies, and customer obligations needs some variation. It might sensibly automate most resource decisions for thousands of small applications while keeping a close planning relationship with the few customers whose requirements can determine the next major capacity purchase. A single process for both would either bury the small applications in planning meetings or leave the largest customers to an automated rule telling them that no capacity is available.

A common idea in cloud cost management is that every team should see its own utilization, capacity, and cost data, and many organizations treat that openness as a differentiator. It is valuable, but only when the people receiving the data can do something with it. If a team has no control over the shared platform’s architecture or its commercial arrangements, a detailed cost report just gives them something else to worry about, while the platform team may be able to make an improvement across thousands of applications that none of those users could make individually. Central management can be the right choice, and I want us to be able to say that without treating it as a failure of participation.

We need the same flexibility in how we understand efficiency. A measured, stateless workload may respond well to a fairly simple CPU-oriented calculation. Apply that calculation to a database without understanding its data, its recovery requirements, or the maintenance it must complete, and you can make a very confident mistake by asking people to shutdown a critical database based on bad assumptions about utilization metrics. An AI system waiting to rebuild an index might look idle until the moment it needs substantial capacity to add a critical item to search results, and a delay of even a few minutes could mean the difference between meeting market demands and missing them entirely. The interesting work is understanding why these cases differ and what evidence would justify a change in each.

I am an ally of the FinOps community, and this book owes it a real debt. J. R. Storment and Mike Fuller’s Cloud FinOps gave this work a shared language, and the FinOps Foundation’s framework connects business, engineering, and finance through a flexible set of capabilities.2 I see FinOps as an essential part of a larger function, one that treats capacity planning and cost management as a single discipline. I am not going to introduce the concept of cloud cost, and this book is less an introduction to the area than a way to put cloud cost management into a conversation about cloud cost and capacity management. What a system costs is hard to separate from how much of it you buy, when you buy it, and what you have promised to keep running when part of it fails.

The SRE and capacity-planning literature is full of hard-won experience, and I draw on it throughout. But much of it approaches capacity and cost through a single lens, and that lens is reliability. Given where that literature comes from, this makes sense, and reliability is, of course, the most important thing in the world. So is launching on time, according to the business. So is staying within budget, according to Finance, and so is shipping the new feature, according to engineering. Everyone in a capacity planning meeting arrives certain that they own the most important part of the conversation, and they are all partly right. The job is to weigh those claims against one another when they depend on the same supply, not to pick one and optimize for it.

The examples, and what to do with them

Our teaching portfolio includes a web service, data platforms, a small AI program, a shared platform supporting a long tail of applications, and a migration with old and new environments running together. These are constructed examples. Their figures are chosen so that we can follow the arithmetic and see how a decision in one part of the portfolio affects another.

The web service gives us a calculation we can understand without spending half the book on its implementation. Then the databases and AI work complicate it, because copying that model would conceal the very constraints we need to manage. The long tail shows how a small change repeated across many consumers can upset a plan built around the largest projects. The migration gives us a necessary improvement whose efficiency numbers get worse before they get better. None of these examples requires an unusually incompetent team. They require people doing reasonable work under conditions that keep changing.

The companion code lets you reproduce selected calculations and follow a decision from observations through forecasting, demand intelligence, supply scenarios, and review. Change the assumptions and see what happens. It is a teaching model, with explicit limits; bringing it into an operating environment would require the evidence, integration, and specialist review that the examples discuss.

My own mistakes belong in the discussion, and I have several decades of mistakes to draw from. A book about judgment would be less useful if I only showed up after the difficult part to explain what everyone else should have done.

Part I, The Decisions, begins with what we are counting, what we are promising, and who can change the promise. Part II, Understanding the Work, develops an operating cycle and gets into measurement, cloud tools, AI, databases, and value. Part III, Demand and Uncertainty, puts the forecasting mathematics alongside customer intelligence and the behavior of a portfolio. Part IV, Running the Practice, follows those decisions through relationships, reserves, wider failures, commitments, and the capacity review.

You may need the simplest practice and the most complicated negotiation on the same afternoon. Keep enough structure to explain the decision, and enough curiosity to notice when that structure has stopped describing the problem. My first manager in this area was correct: capacity and cost is an unwinnable game. I hope that after reading this book you will understand how to change the program.

1 Star Trek II: The Wrath of Khan (Paramount Pictures, 1982), training exercise and Kirk’s explanation; see also StarTrek.com, “Captain Kirk’s Wisest Quotes”, March 22, 2023, and the official film excerpt. The application to capacity management is the author’s interpretation. The manager’s remark is the author’s recollection.

2 J. R. Storment and Mike Fuller, Cloud FinOps, 2nd ed. (O’Reilly, 2023); FinOps Foundation, “FinOps Framework”, accessed September 23, 2026. The framework identifies collaboration, business value, and a flexible set of capabilities; the judgments about selective engagement in this book are the author’s.