● AI Agent Boom
Core Summary of the AGI Debate Sparked by GPT-6 Astra: AI Agents Have Started Taking Over an Entire “Task”
The core point of this issue is not simply that “AI has become smarter.”
The model introduced as OpenAI’s GPT-6 Astra has drawn attention because it has taken a step forward in its ability to read screens, operate programs, verify results, and revise again when it fails.
In particular, the reported 92.7% screen-click accuracy, 72.6% success rate in real computer tasks, and 29 hours of autonomous cybersecurity vulnerability exploration can be seen as signals that AI agents are moving to the center of enterprise productivity innovation.
In this article, we will summarize in news format why Astra has reignited the AGI debate, what tasks it may replace or assist with in real industrial settings, and how it could affect corporate investment and AI semiconductor demand from a global economic outlook perspective.
1. Astra’s Core Change: “AI’s Ability to Directly Handle Computers” Has Changed
The first change to look at in Astra is its computer-use capability.
So-called “computer use” means AI can look at the screen like a person and manipulate the mouse and keyboard, and that ability has improved significantly.
For example, when a user says, “Organize this data in Excel,” the AI opens Excel, finds the menu, enters formulas, and checks whether the results are correct.
This feature itself is not a completely new concept.
Existing AI agents could already open websites, click buttons, or handle simple document tasks to some extent.
But the problem was accuracy.
If AI misread the screen, missed a button’s location, or clicked the wrong spot, the task would get derailed midstream and a human would have to step in again.
Astra’s differentiator is precisely this point.
It is not just that it “can use a computer”; it has become better at understanding the screen more accurately and linking that judgment to the next action.
2. 92.7% Screen Recognition Accuracy: Why This Number Matters
The key benchmark mentioned in the original text is the ScreenSpot Pro test.
This test shows the AI a professional software screen and evaluates whether it can find the correct location when instructed, “Find this button” or “Click here.”
GPT-5.6 Sol, introduced as the previous model, reportedly scored 76.9%.
By contrast, Astra was introduced as reaching 92.7%.
This difference is not just a numerical improvement.
For AI agents to be truly useful in enterprise work, they must look at the screen, understand its meaning, and act precisely.
At the 76.9% level, a person still has to stand by and catch errors continuously.
But when it rises to 92.7%, the number of human interventions falls, and the economics of task automation can change dramatically.
This matters because AI is now starting to move beyond text responses and into real software work environments.
3. PCB Design Case: AI Created Not a Picture, but a Production-Ready Design
Astra’s first industrial case is printed circuit board, or PCB, design.
When a user provided an electronic circuit diagram, Astra opened KiCad, a specialized design program.
It then placed components, connected them with copper traces, and completed a PCB layout that could be used for actual production.
The important point here is that Astra did not simply create an image that looked like a PCB.
The key point is that it performed the work inside the professional software used by real engineers.
PCB layout is a time-consuming task when done manually.
That is because it requires considering component spacing, routing paths, electrical connections, and manufacturability.
If Astra can quickly produce a draft of this work, engineers can spend more time on review and optimization rather than repetitive tasks.
This is a fairly important shift in manufacturing AI adoption and corporate investment strategy.
4. 3D Modeling Case: A Workflow from Blender to Unreal Engine
The area where Astra’s capability becomes more intuitive is 3D production.
According to the original text, Astra used Blender to model a house and then moved the result into Unreal Engine 5 to create a three-dimensional space that a person could walk through.
The point to note in this case is that it did not handle just one program.
It continued the work by moving between multiple specialized tools, building the model in Blender, transferring it to Unreal Engine, and implementing it as if it were a real space.
Traditional generative AI often provides outputs as text or image files.
But Astra’s case differs in that AI enters an actual industrial software pipeline and connects the workflow.
In the gaming, architecture, simulation, and digital twin markets, this kind of capability is likely to lead to productivity innovation.
5. Game Development Case: It Does Not Just Write Code, It Plays and Revises Directly
The game development case shows why Astra is different from conventional coding AI.
Game company PlayCon was introduced as using Astra connected to game engines such as Unity and Godot.
Astra does not merely write code.
It modifies game scenes, runs the actual game, and plays it directly to check whether the changes work properly.
When a problem arises, it fixes it again.
According to the original text, PlayCon had Astra create three game prototypes with different themes based on a graybox stage before graphics were added.
As a result, Astra produced all three prototypes at once, and it was explained that manual correction work by engineers fell by about 50% compared with using GPT-5.6, the previous model.
This matters because game development is not simply a coding problem.
It requires checking whether a character passes through walls, whether buttons are pressed properly, whether objects are placed in strange positions, and whether the gameplay flow feels natural.
In other words, for AI to truly take over work, it must do more than “write code”; it must also handle “screen understanding,” “spatial reasoning,” “execution verification,” and “error correction.”
Astra is significant because it has moved closer to this workflow.
6. Long-Term Autonomous Work Capability: AI Held Onto One Goal for More Than a Day
Astra’s second core point is its long-term task capability.
When assigning work to AI, what matters more than “how smart it is” is often “how long it can keep going without forgetting the goal.”
No matter how excellent an employee is, it is hard to hand over an entire task if they ask every five minutes, “What should I do now?”
Existing AI agents could also carry out work through multiple steps.
But when the task grew long, problems emerged.
They often forgot the original goal, stopped because of a single error, or repeated a method that had already failed.
So in long-term work capability, what matters is not simply how long the AI stays on.
It is the AI’s ability to judge on its own, correct failures, and continue toward the goal before a person steps in again.
7. 72.6% on OSWorld Test: Real Computer Task Success Rate Improved
According to the original text, Astra recorded 72.6% on the OSWorld benchmark, which measures performance in real computer environments by carrying out a task through multiple steps.
GPT-5.6 Sol was introduced as scoring 65.7%.
Numerically, that is an improvement of about 6.9 percentage points.
At first glance, that may not seem large.
But in actual workflow automation, this difference is quite significant.
Whether an AI agent fails three out of ten times or two out of ten times directly affects operating costs and workforce allocation.
In areas such as repetitive work, data organization, QA testing, security checks, and design assistance, even a slight increase in success rate expands the range of tasks that can be automated.
8. 3D Creation of Fallingwater: A Long Project Lasted 29 Hours
YouTuber Pat Simmons had Astra create a 3D version of Fallingwater, the famous building by Frank Lloyd Wright.
Astra searched for drawings and related materials, operated Blender, checked the output via screenshots, and repeated the revision process.
It was introduced as taking more than 29 hours to complete the final result.
However, interpreting this case as “AI worked entirely on its own” is a bit of an exaggeration.
That is because there was human approval and intervention along the way.
Still, the meaning is clear.
It shows that AI can now continue a single complex task for tens of hours.
This is a highly important change in the fields of architectural modeling, virtual space creation, the metaverse, and digital twins.
9. Manhattan Implementation Project: Manager AI and Execution AI Worked Together
In another case, Matthew Schumer, an American venture entrepreneur and AI developer, used Astra to carry out a project in Unreal Engine that recreated Manhattan down to its buildings and streets.
The project was said to have taken place over the course of a week.
However, it was not a case of leaving one Astra instance running continuously for a week.
Schumer used a “manager loop” approach to continue tasks through multiple stages.
In simple terms, one AI acts as the team leader for the overall project.
It breaks tasks into stages and determines the next thing to do.
Another AI handles the actual production.
When the execution AI finishes one stage, the manager AI hands off the next task again.
This structure is likely to become a very important method in future enterprise AI agent operations.
That is because a system in which multiple AIs divide roles and create a management structure is more realistic than one super AI doing everything.
10. Cybersecurity Case: 29 Hours of Exploration with Minimal Human Intervention
The strongest case is the cybersecurity evaluation.
OpenAI reportedly conducted a cybersecurity assessment before launch to confirm Astra’s long-term autonomous work capability.
Researchers provided the initial objective, source code, and basic analysis tools.
However, they did not give step-by-step instructions such as “Run this next” or “Try again using this method.”
The goal was to let Astra explore directly with minimal human intervention.
As a result, it reportedly took about 29 hours to find an unknown vulnerability in the browser and create a working attack method.
This case carries very sensitive implications for the AI industry.
On one hand, it can rapidly identify security vulnerabilities and improve defenses.
On the other hand, it can also increase the potential for misuse.
That is why AI regulation, cybersecurity investment, and corporate security budgets are all likely to grow together in the future.
11. Changes in Internal Reasoning: Possibility of a Recursive Depth Approach
The original text also cited reporting from The Information, a U.S. IT media outlet, saying that Astra uses a “recursive depth” approach in a limited way.
Recursive depth can be described as a method of increasing the depth of internal computation by passing through the same computational layer multiple times.
If the conventional method is a structure that passes once through computation layers A, B, and C to produce an answer, recursive depth is closer to going back through the same layers to think more deeply.
However, OpenAI has not officially confirmed that this technology is the direct reason for Astra’s long-term work capability.
Therefore, it is safest to view this part as closer to industry reporting and speculation than as confirmed fact.
12. The Real Bottleneck in Long Tasks: Memory and Work Log Management
For AI to work for hours at a time, better reasoning alone is not enough.
It must remember what it did a few hours ago, which methods failed, and which files were modified.
Existing AI agents used a method of summarizing and compressing earlier work when the task record became long.
The problem is that important details can be lost in the summarization process.
For example, if the information about why a certain method failed earlier disappears, the AI may repeat the same mistake hours later.
According to the original text, OpenAI configured Astra in Codex so that when it continues a long task, it leaves separate work notes and can search previous conversations or tool-use records if needed.
It also does not always stop even when it needs to ask the user something during the task.
It continues tasks that can proceed regardless of the answer, and only waits for the user when a decision is needed that will significantly affect the result.
This change is extremely important for real-world enterprise adoption.
If AI stops because of a single question, the automation effect is reduced.
13. Why the AGI Debate Has Grown Again
Astra has reignited the AGI debate because it combines two capabilities.
First, the ability to directly see and manipulate computers.
Second, the ability to maintain a single goal for a long time while repeating failure and revision.
Even if AI can operate a computer well, if it quickly forgets the goal, it cannot handle complex work.
Conversely, even if it can think for a long time, if it cannot actually handle real programs, a person still has to step in halfway through.
Astra’s case shows that these two abilities are advancing at the same time.
That is why some people are saying it is “close to AGI.”
However, there is also something to be careful about here.
There is still no clear consensus definition of AGI.
Even if Astra can perform a wide range of industrial tasks, it would be hard to say it has already reached human-level general intelligence.
Rather, at this stage, it is more realistic to say it has become closer to an AI agent that can take over an entire task.
14. The Most Important Change for Businesses: Operational Data Matters More Than Passing a POC
At the beginning of the original text, there was also mention of the Snowflake World Tour Seoul recap live seminar.
The important message there is the question: “We thought passing the POC meant everything was done, so why does it not work for our company?”
This is a problem that AI adoption companies actually face most often.
It works well in demos.
But once connected to internal company data, permission systems, security policies, and existing business systems, the difficulty suddenly increases.
If the data immediately available to the AI agent is insufficient, if data quality is poor, or if access permissions are complicated, performance drops sharply.
In other words, future corporate competitiveness will not end with simply using a good AI model.
What becomes core is how well a company has organized its own data and how it has created a work environment that AI can safely access.
This is one of the most important points in enterprise investment in the AI era.
15. Global Economic Outlook: AI Agents Change Software Budgets Before the Labor Market
If long-term autonomous AI agents like Astra spread, the first thing to change is likely not the overall labor market but the structure of corporate software budgets.
Rather than replacing people immediately, companies are likely to start by attaching AI agents to the work tools already used by employees.
For example, a designer can hand over PCB drafts to AI and focus on review.
A game developer can hand over repetitive testing and prototype production to AI.
A security team can delegate part of vulnerability exploration to AI and focus on prioritization and response strategy.
This change can create productivity innovation, but at the same time it can also further stimulate investment in AI semiconductors, cloud infrastructure, cybersecurity, and data platforms.
In particular, the fact that AI works for 29 hours or on a weekly basis means inference costs continue to accumulate.
The smarter models become, the more computational resources and cost management companies will need.
16. The Most Important Point Easy to Miss in Other News
The real significance of this Astra issue is not the sensational phrase “AI worked for more than a day.”
The core point is that AI is becoming a new unit of labor.
Until now, companies assigned tasks to people, and software was a tool people used.
But as AI agents like Astra develop, that structure changes.
People set the goals and criteria, and AI directly operates multiple software tools to produce results.
In other words, software shifts from being a human tool to being an AI workplace.
Once this change takes hold, corporate competitiveness will be divided into three areas.
First, how well the data accessible to AI has been organized.
Second, whether there is a sandbox environment where AI can safely be tested even when it makes mistakes.
Third, whether there is an audit system that can track and verify AI’s work process.
Many news reports focus on model performance numbers.
But in actual corporate settings, data permissions, security policies, verification processes, and cost management determine the outcome.
Going forward, the companies that succeed in adopting AI agents will likely not be “companies that used a good model,” but “companies that built an operating system where AI can work.”
17. Summary by Industry
Manufacturing and electronic design: AI-assisted design may spread in tasks like PCB layout, where repetition and expertise are both required.
Game development: Automation may become possible across coding, scene editing, play testing, and bug fixing.
Architecture and 3D modeling: The speed of turning drawings into 3D implementations, recreating buildings, and creating virtual spaces may increase.
Cybersecurity: Vulnerability discovery and attack simulation automation will strengthen, and both defensive capabilities and misuse risks may rise at the same time.
Enterprise AI platforms: Internal company data, access control, workflow automation, and cost optimization may emerge as core infrastructure.
18. Investment Checkpoints to Watch
First, as AI agents are used more widely, cloud inference costs and AI semiconductor demand are likely to remain important.
Second, companies are likely to allocate budgets to agentic AI that manipulates actual business software rather than simple chatbots.
Third, as long-term autonomous work increases, cybersecurity and monitoring solutions become more important.
Fourth, data platform companies may be evaluated as key infrastructure that provides well-organized data usable by AI.
Fifth, from a global economic outlook perspective, AI is likely to contribute first to higher throughput and faster development rather than immediate cost reduction.
19. Conclusion: Astra Is Closer to “Working AI” Than a “Smarter Chatbot”
The significance of Astra goes beyond ChatGPT simply giving better answers.
It is developing in the direction of seeing the screen, manipulating programs, verifying results, revising when it fails, and sustaining goals over long periods.
When this combination enters actual industrial settings, AI becomes not just a support tool but an executing entity within the workflow.
It is still too early to say whether it is AGI.
But for companies, a more practical question than the AGI debate is important.
“How much of our company’s work can we hand over to AI in one piece?”
“How will we verify and control AI when it makes mistakes?”
“Are our data and access structures ready for AI to work?”
The companies that can answer these questions are likely to get ahead in the next AI competition.
< Summary >
The key point of the model introduced as GPT-6 Astra is screen recognition and long-term autonomous work capability.
It was reported to have recorded 92.7% accuracy in screen manipulation on the ScreenSpot Pro test and 72.6% success in actual computer tasks on the OSWorld test.
Cases close to real industrial work appeared, including PCB design, 3D modeling, game development, building implementation, and cybersecurity vulnerability exploration.
In particular, the case of exploring vulnerabilities for 29 hours shows that AI agents can maintain a goal and work for long periods of time.
The real core point is not that AI became smarter, but that it is evolving into an execution-type AI that directly manipulates software and carries out work units.
From now on, corporate competitiveness will depend less on model selection and more on how well data organization, access control, security systems, and verification processes are in place.
[Related Articles…]
How AI Agents Are Reshaping Enterprise Productivity Strategy
AI Semiconductor Investment Cycle and the Global Economic Outlook
*Source: [ 티타임즈TV ]
– AI가 하루 넘게 혼자 알아서 일했다!


