● AI Prompts Now Cost Too Much
Core Takeaways from Astra-Era Prompt Engineering: You should make AI “waste less,” not “work more”
The core point of this piece is quite clear.
High-performance AI agents are no longer a problem because they fail to work; they are a problem because they do too much.
Especially with models like Astra that automatically search, read files, call skills, and even perform verification, a single line of prompt directly connects to token costs, API costs, work productivity, and enterprise AI investment efficiency.
In the past, writing things like “always check,” “read every time,” “think at length,” or “explain step by step in detail” seemed like careful instructions, but in practice, they can become cost bombs that keep triggering unnecessary file access, skills, searches, and verification.
In this article, we will organize the differences between bad prompts and good prompts, the structure by which Astra-type AI agents burn tokens, the prompt-design methods enterprises must change to improve AI productivity, and the truly important points that other news coverage often misses.
1. The news core takeaway: the era of prompts has changed
In the past, the core of prompt engineering was “explaining as much as possible so the AI does not make mistakes.”
Because models were weaker, it was effective to give a lot of background, break tasks into steps, make them review repeatedly, and instruct them to think as long as possible.
But with the latest AI agent models like Astra, the situation has reversed.
Because the model investigates, verifies, uses tools, and calls skills on its own, overwork happens unless the user limits the scope.
In other words, what matters now is not “how do we make AI move more,” but “how do we decide how far AI should go and where it should stop?”
- Existing approach: explain more and make it review more
- Astra-era approach: let it read only the necessary materials, do only the necessary verification, and stop at the designated point
- Enterprise perspective: a prompt is not just a sentence, but a cost-control mechanism and an AI governance design document
This change is also important in global economic outlook discussions.
As companies adopt AI agents, the first issue they run into is token cost and operational efficiency, even before productivity gains.
As the scale of AI investment grows, prompt-design ability is becoming not just a personal skill, but a core capability that determines enterprise digital transformation costs.
2. The worst prompts: expressions like “every time,” “always,” and “all”
The expressions most strongly criticized in the source were words like “every time,” “always,” “all,” and “everything.”
On the surface they look like careful instructions, but for an AI agent they can be interpreted as meaning that it should reread unnecessary materials and review unrelated skills on every task.
Bad example 1
Before every edit, read architecture.md, database.md, and deployment.md.
The reason this prompt is dangerous is simple.
Even if the edit is light work like changing the color of a frontend button, it may cause the agent to read all of the architecture, database, and deployment documents.
During this process, input tokens keep increasing, and the model performs unnecessary verification as well.
Good example 1
Only refer to architecture.md when changing service boundaries.Only refer to database.md when there are schema changes.Only refer to deployment.md during deployment preparation.Do not read the above documents for any other edits.
The core point of a good prompt is clearly defining “when to read” and “when not to read.”
That is how the AI agent selectively loads only the necessary documents and reduces token costs.
Bad example 2
Create and verify PostgreSQL schema migrations.Use database query models or tasks related to data persistence.
This sentence is too broad because the phrase “tasks related to data persistence” is vague.
The model may try to review every database skill and document that could possibly be related.
Good example 2
Use this guidance only when adding or changing migrations.Check related queries and schema impact only during the rollout review stage.Do not apply this guidance to simple read-query edits.
Good prompts include the timing of application, exclusion conditions, and stopping conditions together.
Without those three, AI will begin over-verifying in the name of “being safe.”
3. The structure by which Astra-type AI burns tokens
With agent models like Astra, tokens are not consumed only in the user’s question and answer.
Costs also arise while reading files, running searches, calling skills, leaving verification logs, and performing internal reasoning.
- Input tokens: the prompt, files, documents, codebase, and search results provided by the user
- Output tokens: the answer, explanation, logs, and reports the model shows the user
- Reasoning tokens: the process the model internally thinks through, even if the user cannot see it
- Tool execution costs: search, file reading, code execution, deployment, rendering verification, and similar actions
- Skill invocation costs: costs that occur when skills are called even though their necessity is unclear
A particularly important point is that “invisible reasoning can still be a cost.”
Even if the user instructs, “Don’t think, just answer,” the model does not actually stop internal reasoning completely.
So if you want to save tokens, the answer is not “don’t think,” but “limit the result scope and the verification scope.”
Inefficient instruction
Save tokens by not thinking and just answering.
Efficient instruction
Write the final answer in 5 sentences or fewer.Provide only 2 key pieces of evidence.Do not perform additional verification, and answer only within the materials currently provided.
4. “Think at length” can now become a cost bomb
Many users still write things like “think step by step,” “consider it at length,” or “carefully review all issues.”
With older models, these expressions helped improve answer quality.
But with Astra-type models, which already actively perform verification and tool use, such instructions can lead to overwork.
Bad example
Think through all issues step by step and explain them at length.Do not miss even small possibilities, and repeatedly review multiple alternatives.
This kind of prompt is similar to telling the model, “Let’s burn all the tokens today.”
Not only will the answer become longer, but unnecessary verification and alternative generation are also likely to occur.
Good example
Go through only the important decision-making process and provide only 3 key pieces of evidence.Compare only the 2 most realistic alternatives.Do not guess at areas with high uncertainty; only mark them.
The core point is not to block the model’s thinking, but to define the size of the deliverable and the level of verification the user actually needs.
5. More skills are not better; skill descriptions must be accurate
An especially important part of the source is skill design.
Many users keep adding skills to make AI agents more useful.
They stack features such as PDF summarization skills, document processing skills, contract review skills, data automation skills, and code analysis skills.
The problem is that if the skill name and description are vague, the model will invoke the skill too often.
For example, if there is a skill called “document processing,” it may be called for almost every document task such as PDF summaries, legal contract review, report analysis, and email organization.
Vague skill names
Document processingPDF summarizationData analysisWork automation
Good skill names
Legal contract risk reviewFinancial statement key metric summaryPostgreSQL migration impact analysisYouTube performance dashboard generation
The more specific the skill name and description are, the fewer unnecessary invocations occur.
The ideal way for an AI agent to work is to first look at minimal metadata such as the skill name and description, and only read detailed files or scripts when necessary.
The source described this as progressive disclosure, meaning a structure that opens only the information needed at each step.
This is very important from a corporate perspective.
When building an enterprise AI platform, what matters more than creating lots of skills is designing the skill taxonomy, invocation conditions, and exclusion conditions.
If that is not done properly, the AI agent may increase cloud costs and API costs rather than improving work productivity.
6. “Read only when truly necessary” may not be a good instruction either
An interesting point is that the phrase “read only when truly necessary” is not necessarily a good prompt either.
It sounds reasonable to people, but for the model the standard for “truly necessary” is vague.
It may decide something is necessary and read too much, or on the other hand, fail to read something even when it is needed.
Therefore, the standard must be written specifically.
Vague instruction
Read related documents only when truly necessary.
Specific instruction
Only read api-spec.md when changing the API response structure.Only read schema.md when columns are added, removed, or their types change in the database.Do not read the above documents when only changing UI text or styles.
Good prompts define both what should be done and what should not be done.
Only with this structure can the model reduce overwork while keeping AI productivity and lowering costs.
7. Astra’s characteristic: show less of the thought process, and win through results and actions
The source mentioned an important Astra characteristic: it does not directly show chain-of-thought, or CoT.
Simply put, older models showed step-by-step reasoning like writing out the process for solving a math problem by hand, while Astra-type models are closer to doing it internally like mental calculation and then producing the final result.
The advantage of this approach is efficiency.
It can repeatedly calculate only the necessary parts, skip easy parts quickly, and still reach the answer.
The source evaluated this as feeling strongly like hallucination is reduced.
But there is also a downside.
Because users cannot easily see the process by which the model reached its conclusion, watching only the explanation is not enough.
You must also confirm which tools were actually executed, which files were read, and what actions were taken.
This is especially connected to security and compliance issues in enterprise AI environments.
Logging and permission management become important for what data the AI agent accessed and what tasks it performed.
8. Verification should be done not “a lot,” but only “as much as specified”
Astra-type models are good at verification.
The problem is that they try too hard to be good at it.
If the user does not specify the scope of verification, the model may capture screenshots, check rendering, inspect mobile screens, check build errors, and review links and buttons in full.
This difference also appeared in website replication experiments.
The source introduced an experiment in which an interactive webpage like Anthropic’s research page was translated into Korean and implemented with similar design and scroll interactions.
Within the specified verification items, Astra showed fairly strong reproducibility for the original screen, scroll behavior, buttons, mobile layout, Korean line breaks, and build errors.
The core point here is that the verification items were limited to about 6.
Perform only the 6 verification items below.1. Visual layout consistency2. Scroll animation behavior3. Button, link, and menu behavior4. Desktop layout5. Mobile layout6. Korean line breaks and build errors
If it had only said “verify carefully,” it likely would have performed far more verification, increasing both tokens and time.
The better the model is at verification, the more the verification scope must be limited.
9. Low, medium, high, ultra: reasoning intensity must be used according to task difficulty
One practically useful point in the source is the criteria for selecting reasoning intensity.
Because Astra can produce fairly good results even in low reasoning mode, using a high-performance mode from the start may be inefficient from a cost perspective.
| Reasoning intensity | Recommended use cases | Notes |
|---|---|---|
| Low | General work, simple analysis, light code edits, short document summaries | For most tasks, starting here is the most economical choice. |
| Medium | Reports combining multiple documents, evidence-based analysis, dashboard generation, complex tasks | The verification scope and output format must be specified. |
| High | Code problems with complex constraints, repeatedly failed tasks, structural design | Costs may increase, so it should be used after organizing failure patterns. |
| Ultra | Long-term projects, complex reasoning, high-difficulty automation, multi-agent tasks | Because autonomy and persistence are strong, overwork can easily occur. |
In particular, ultra mode was described as a “persistent and tenacious model.”
If the user does not define the goal and stopping conditions clearly, the model may create a main agent and sub-agents, repeat image rendering, and consume enormous amounts of tokens.
This is a very important point from the standpoint of enterprise AI cost management.
Using a high-performance model does not always produce a better ROI.
Selecting the right model and reasoning intensity for the task difficulty determines AI investment efficiency.
10. Practical prompt structure: writing from the output first is effective
One good method that repeatedly appeared in the source was to write the deliverable or goal first.
In the past, prompts often began with role assignment, background explanation, and then instructions, but for Astra-type models, placing the goal and deliverable first was introduced as working more clearly.
Recommended structure
[Deliverable]Write first what needs to be created.[Input materials]Write the URL, files, and data scope to use.[Tools to use]Limit the plugins, search, and data analysis tools.[Conditions]Write the design, language, format, length, and analysis criteria.[Prohibitions]Write exclusion ranges such as no guessing, no number manipulation, and no unnecessary sentiment analysis.[Verification scope]Limit the verification items by number.[Stopping conditions]Define how far completion should go before stopping.
For example, if you are building a YouTube channel dashboard, something like this is good.
The deliverable is a YouTube playlist performance dashboard.Collect only video-level data from the provided URL.Use only the data plugin and the site semantic layer.Visualize views, video length, monthly upload count, cumulative views, and median values.Do not perform comment sentiment analysis, causal inference, or arbitrary number generation.After completion, provide only 5 key insights and the dashboard link.
This method is effective at preventing unnecessary analysis.
It can stop the model from expanding “make a good dashboard” into comment sentiment analysis, causal inference, and confidence evaluation.
11. Multi-turn conversations still require caution
Even with the latest models, results can become fuzzy as multi-turn conversations get longer.
The source noted that Astra felt more stable in multi-turn interaction than existing models, but also explained that there is a research trend suggesting quality can start to wobble after 7 turns or more.
In practice, it is good to set these rules.
- If the answer suddenly becomes strange, start again in a new window.
- When moving to a new window, transfer the core state in a structured form, not as a simple summary.
- Separate and organize files, decisions, current issues, completed tasks, and remaining tasks.
- For long-term work, operate a separate memory file such as memory_summary.md.
However, even the word “summary” should be handled carefully.
If you simply shorten it, important context may be lost.
A good summary is compression that preserves necessary terminology and decisions without damaging the original meaning.
Good memory organization example
Current goal:Completed work:Important decisions:Files used:Remaining issues:Failed patterns that should not be repeated:Conditions that must be maintained in the next task:
12. The most important point that other news does not explain well
First, prompts are now not a writing skill but a cost-control skill.
Many news articles focus on comparing AI model performance.
But in real work, what matters more than how smart the model is is how economically that intelligence is used.
When you include token costs, cloud costs, API usage, and verification time, a prompt is effectively a corporate cost-management document.
Second, the skill taxonomy becomes a new competitive advantage.
In the AI agent era, the companies with the most skills are not the ones with the advantage.
The companies that build skill systems with accurate names, invocation conditions, exclusion conditions, and priorities will have the advantage.
This part is likely to become an important differentiator in the enterprise AI market.
Third, an AI that verifies “a lot” is not always a better AI.
Verification improves quality, but it also raises costs.
Therefore, not defining verification items is similar to outsourcing work with no budget.
Enterprises should grade verification levels according to business risk.
Fourth, the cause of AI agent adoption failure may be the data and operational structure, not the model.
The message from the Snowflake World Tour Seoul recap seminar mentioned at the beginning of the source connects to this as well.
You may think that simply attaching an agent means it will find and use data on its own, but in real enterprises the data access permissions, quality, semantic layer, and tool connections are all different.
Passing a POC does not mean it will succeed immediately in the production environment.
Fifth, the more powerful the model, the smarter the user must be.
Ironically, the smarter the model becomes, the more precisely the user must restrict it.
You should not say “just do it for me,” but instead instruct: “Within this scope, using only this material, do only this level of verification, and stop here.”
This capability is likely to become a core work skill in the AI era.
13. Astra prompt checklist for enterprise practitioners
- First remove expressions like “every time,” “always,” “all,” and “everything.”
- Clearly specify when files should be read.
- Also write the conditions under which files should not be read.
- Do not make skill names overly broad; make them specific around the business purpose.
- Limit the answer length by sentence count, number of tables, or number of items.
- Limit verification items by number.
- Limit the search scope to authoritative sites, specific media outlets, or specific data sources.
- Receive only a core summary, not the full tool execution log.
- For long-term work, keep only checkpoints and prevent unnecessary log generation.
- Start in Low mode, and raise it to Medium or High only when there is a failure.
- Use Ultra only for complex long-term work or high-difficulty reasoning.
- For important work, use a manual approval structure rather than automatic approval.
- However, if approvals are requested too often, productivity drops, so approval criteria must also be set.
14. A good prompt template you can use right away
Deliverable:[Write the desired final result in one sentence]Input materials:[Files, URL, and data scope to use][Do not use any other materials]Task scope:[3 to 5 things to do]Exclusion scope:[Things not to do][No guessing, no arbitrary generation, no unnecessary searching, etc.]File usage criteria:[In what situations to read which files][Otherwise, do not read them]Skill usage criteria:[Skill name to use][Invocation conditions][Conditions under which it should not be invoked]Verification criteria:[Limit verification items by number]Output format:[Table, list, report, code, etc.][Length limit]Stopping conditions:[How far completion should go before stopping][Information to provide after completion]
The core point of this template is to give the model freedom while clearly locking down the areas where costs can grow.
In particular, it is a good idea to include file usage criteria, skill usage criteria, verification criteria, and stopping conditions.
15. Conclusion from an AI Trend perspective
The arrival of Astra-type models is not simply the appearance of a smarter chatbot.
It is closer to meaning that AI agents have begun to behave like work performers.
So prompt engineering going forward is likely to evolve not as the skill of writing pretty sentences, but in a direction that handles work design, cost control, risk management, and data-access policy together.
From the perspective of enterprise productivity, the opportunity is large.
Tasks such as dashboard creation, document analysis, code editing, webpage implementation, and report writing can be handled much faster.
But at the same time, if not managed, token costs can rise quickly, and operational efficiency can drop due to unnecessary verification and tool execution.
In the end, the person who uses AI well in the future is not the one who makes AI do more.
It is the one who accurately designs not only what AI should do, but also what it should not do, based on the assumption that AI is already capable enough.
< Summary >
In the era of Astra-type AI agents, expressions like “every time,” “always,” and “all” can significantly increase token costs.
A good prompt clearly defines not only what should be done, but also what should not be done, when files should be read, the verification scope, and the stopping conditions.
With skills, it is more important to design the name, description, and invocation conditions accurately than to create many of them.
More powerful models need prompts that prevent overwork.
From a corporate perspective, prompt engineering is becoming a core capability that determines AI productivity, API costs, and digital transformation ROI.
[Related Articles…]
*Source: [ 티타임즈TV ]
– 토큰 까먹는 나쁜 프롬프트 예시, 좋은 프롬프트 예시 (강수진 박사)


