Prompt format barely changes AI accuracy, but can raise costs 45%

Jul. 29, 2026
By AI, Created 04:53 UTC, Jul 29, 2026, AGP -

A July 2026 benchmark by Paulo Teixeira tested seven ways to write the same AI instruction across 27 models and 17,850 graded answers. The study found prompt format made little difference to accuracy, but token usage — and therefore cost — varied sharply.

Why it matters: - The benchmark suggests prompt style is often a cost decision, not an accuracy decision. - For teams sending instructions at scale, a 45% token gap can become a recurring bill, even when answer quality stays the same. - The study also shows that model settings and execution context can matter more than the prompt format itself.

What happened: - Paulo Teixeira benchmarked seven prompt formats across 27 AI models, 34 complete runs, and 17,850 graded answers. - The test concluded in July 2026 and covered six providers, two languages, and seven prompt formats. - The benchmark measured exact-match accuracy and token usage for each prompt style. - Six of the seven formats landed within three-quarters of a percentage point of one another on accuracy. - Token use ranged from 473 tokens per prompt for Markdown in English to 686 for XML in Portuguese. - The difference in token use translated into a cost gap of about 45%.

The details: - The study used 525 cells per run, for a total API bill of roughly US$900. - Each answer was checked character by character against a code-generated answer key. - No human judge scored the output. - The benchmark included three task families: GRID, CASCADE, and VOLTGRID. - GRID tested visual induction in an ARC-AGI-2 style. - CASCADE chained three transformations, where one early mistake could break the final answer. - VOLTGRID was created for the study and did not exist online before publication. - VOLTGRID left one rule out on purpose, forcing models to infer it from execution traces. - One puzzle had no correct answers in 1,190 attempts. - The easiest puzzle was solved 83.9% of the time. - The highest overall score in the study was 77.5%. - NTC prompt engineering finished with the highest accuracy at 49.29%. - Markdown in English was the cheapest format and also sat on the efficiency frontier. - XML in Portuguese finished at 49.25%, just one correct answer behind NTC across 2,550 cases. - Under McNemar testing, that gap produced p = 1.00, meaning no detectable difference. - Portuguese scored 49.01% overall versus 47.79% for English, a 1.22-point gap across 7,650 cells per language. - Run by run, Portuguese won 16 times, English won 16 times, and two runs tied. - Switching a model's extended reasoning on moved accuracy by 41.7 points. - Changing the runtime environment moved accuracy by as much as 20.6 points. - Changing the framework moved it by a median of 12.0 points within a run. - Changing instruction language moved it least. - Across the whole benchmark, the NTC framework used 507 tokens per prompt, compared with 473 for the leanest format. - In VOLTGRID, NTC was both the most accurate format at 46.00% and the second-leanest at 444 tokens. - XML in Portuguese used 665 tokens in VOLTGRID, making NTC 33.2% cheaper on that task. - Across 93 real production prompt pairs, Teixeira said token count fell from 418,158 to 182,121 after rewriting. - That production sample showed a median reduction of 55.72% per pair and a best case of 73.7%. - One production pair grew 3.3% larger. - Teixeira reported reductions of 70% to 80% on system prompts around 40,000 tokens in live work. - The report says a top-tier model refused to answer 43 of 525 cells because a safety classifier blocked it before the task. - The prompts for six public formats, the rankings, and the full 17,850-cell dataset are public under a licence allowing use and adaptation, including commercial use, with attribution. - The answer keys and the NTC notation itself are withheld.

Between the lines: - The study's main conclusion is that prompt notation is often overvalued compared with context, reasoning settings, and environment. - The pooled average can hide large within-run swings because gains for one model can cancel losses for another. - Teixeira says the average does not measure the effect; it measures what remains after effects cancel out. - The benchmark's paired design gives its claims more weight than an average-only comparison would. - The work also undercuts the common advice to default to English prompts. - The result does not show Portuguese is always better; it shows the best language depends on the model and the task. - NTC's strongest case appears in longer, denser instructions where clarification can also reduce token count. - Teixeira frames the method as clarification first and compression second. - The study also leaves room for future models, since the tasks and answer keys were designed to stay usable against newer systems.

What's next: - Teixeira says one finding from the same dataset will be published separately: the safety-blocked model that skipped 43 cells. - The open dataset allows others to rerun the paired tests and challenge the results. - The benchmark's structure gives future prompt engineering comparisons a public baseline, not just a ranking.

The bottom line: - On this benchmark, prompt format changed the bill far more than the answer. - The bigger levers were whether reasoning was enabled, where the model ran, and how much instruction content had to be sent each time.

Disclaimer: This article was produced by AGP Wire with the assistance of artificial intelligence based on original source content and has been refined to improve clarity, structure, and readability. This content is provided on an “as is” basis. While care has been taken in its preparation, it may contain inaccuracies or omissions, and readers should consult the original source and independently verify key information where appropriate. This content is for informational purposes only and does not constitute legal, financial, investment, or other professional advice.

Sign up for:

Sustainable Planet Portugal

The daily local news briefing you can trust. Every day. Subscribe now.

By signing up, you agree to our Terms & Conditions.

Share this page:

Advanced Search Options

Search for:

Search scope:

Type:

Search in:

Date range:

The last

Sort by:

Sign up for:

Sustainable Planet Portugal

The daily local news briefing you can trust. Every day. Subscribe now.

By signing up, you agree to our Terms & Conditions.