Not Every Section of Your Report Needs Your Most Expensive Model
The Premium-Model Default Nobody Questions
Most teams choose their most capable model once and never revisit the decision. From then on, every section of every report, whether it's a boilerplate methodology note or a nuanced strategic assessment, runs through the same premium model at the same premium token rate.
It's the AI version of over-provisioning cloud infrastructure "just in case." Whether or not something needs the most powerful model on the market, if that’s what the team picked on day one, then that’s what every response will use no matter what.
Model Choice Affects Cost as Well as Quality
When teams pick a model for reporting, the question is almost always which one writes the best output. That treats model selection as a quality dial and skips the second one: cost. Any serious approach to AI model cost optimization has to account for both.
A stronger model charges a premium on every token it processes, including the tokens spent on sections where that strength adds nothing. A findings summary, a methodology description, or a compliance disclaimer doesn't need the reasoning depth a strategic recommendation does.
Where Heavy Models Earn Their Keep
Don’t get us wrong; some sections do benefit from a premium model. Executive summaries have to distill complex, sometimes contradictory source material into a confident narrative. Strategic recommendations need tone-sensitivity and the ability to weigh competing inputs before landing on a defensible position. And cross-source synthesis, where the model reconciles scattered data points into one analytical thread, is the kind of work top-tier models were built for.
In these sections, the extra token spend buys visibly better output, and downgrading the model puts quality at risk.
Where Lighter Models Perform Just as Well
Take recurring reports, for example. Much of the production is scheduled and repeatable: fixed formats, extraction-heavy work, and little room for interpretation. A lighter, faster model handles that work well.
Real-life examples of these report sections include:
Findings summaries that reorganize data points into bullet takeaways follow predictable patterns.
Methodology and boilerplate sections, such as scope statements, sourcing descriptions, and regulatory disclaimers, barely change between cycles.
Structured data pulls, where the model extracts entities, dates, figures, or indicators and slots them into a defined format, are close to mechanical.
In each case, the analytical judgment has already been made, by the analyst who built the template, by the framework the organization adopted, or by the structure of the data itself. The model is assembling and formatting, not interpreting. When the format is defined and ambiguity is low, output from a fast model and a premium one looks much the same, while the token cost stays very different.
What This Looks Like in a Six-Section Report
Take a weekly threat brief with six sections: executive summary, key findings, methodology, an actor activity table, recommendations, and a sources appendix. Two of them, the executive summary and the recommendations, call for a premium model's reasoning and tone. The other four are structured or formatted work.
If you run all six on a top-tier model, then every section is billed at the premium rate. However, if you assign a lighter model to the four structured sections, then only two are. Repeat that every reporting cycle, across every analyst on the team, and the savings show up in token spend while the two premium sections stay exactly the same.
Why Section-Level Assignment Gets Overlooked
Teams that get intentional about AI spend usually do it at the report level: pick a model for the deliverable, then move on. Unfortunately, this workflow produces a blanket allocation, with every section billed at the same per-token rate whether or not the work inside it needs that model.
Two factors contribute to this cycle. The first is habit: workflows built around one model per report harden into standard practice, and nobody revisits them. The second is visibility: without a view of which sections use the most tokens and which a lighter model could handle, there's nothing to act on, and the cost shows up as a single line on the invoice.
Switching a section to a lighter model doesn't have to lower its quality. If a lighter model already handles that section well, the output stays the same and only the cost drops.
How Indago Supports Section-Level Model Selection
In Indago, analysts can actually use different AI models for different sections of the report. It works like this: a report template saves one default model for the full report, and analysts can choose a different model for any individual section when they regenerate it. A lighter, faster model can take a structured findings summary or a methodology section, while a stronger reasoning model handles the executive narrative and the recommendations. Indago has many different models to choose from, too.
Because regeneration works one section at a time, an analyst can revise a single section with targeted instructions, or switch its model, without regenerating the whole report. The sections that already work stay as they are.
The analyst stays in charge throughout. They decide which model fits each section and review every output before it goes anywhere.
Cutting Token Spend Without Cutting Quality
Most conversations about enterprise AI budget management focus on negotiating rates, trimming seats, or capping usage. Section-level model selection takes a different route: it removes waste, meaning premium tokens spent on work a lighter model handles just as well, and leaves everyone's access untouched.
That can reduce AI token spend without interrupting the sections that need a premium model. The executive summary still gets the nuanced reasoning it requires, and the structured findings still come out accurate, because each section runs on the model that fits its complexity. Teams that manage model choice this way can steer their AI spend, and AI reporting cost efficiency improves without losing depth where it counts. The savings can go toward wider coverage, more analysts, or deeper investigations.
Book a demo to see how section-level model selection works in Indago and what it could mean for your team's token spend.