● AI ROI Sinks Without Data Governance
The First Things AI-Savvy Companies Fixed: Snowflake World Tour Seoul Core Summary
If you introduced an AI agent but its results are lower than expected, the problem is more likely the state of your company data than model performance.
The core takeaway of Snowflake World Tour Seoul was very clear.
The real battleground for enterprise AI is not choosing a better generative AI model, but creating data that AI can understand.
The common thread running through the cases from KB Kookmin Bank, Amorepacific, KT, Tossplace, GS Construction, Neowiz, Banksalad, and Nepes is exactly this.
Data governance, semantic views, Apache Iceberg, Cortex agents, and data streams are no longer just IT terms; they are becoming core infrastructure for enterprise productivity and digital transformation.
In particular, this event was less about “How do we introduce AI?” and more about “How should we change our data structure so AI can actually make money?”
1. Core Message of the Event: The Cause of AI Project Failure Lies More in ‘Data’ Than in the ‘Model’
The main theme of Snowflake World Tour Seoul was “Making AI Work for Business.”
Simply put, enterprise AI has now moved beyond the experimental stage and must deliver real business results.
Many companies have introduced chatbots, generative AI, and AI agents, but in practice there are many cases where they fail to produce useful answers.
At that point, people often think, “Is the model lacking?” or “Should we use a better LLM?”
But the message repeated throughout the event was the opposite.
What is needed before a good AI model is a data foundation that AI can trust and read.
- If each department defines sales, customers, subscribers, or performance differently, AI will give different answers to the same question.
- If data is scattered across multiple systems, AI agents cannot grasp the full context.
- If structured and unstructured data are separated, answer quality for business questions drops.
- Without metadata and a business glossary, AI cannot understand what the numbers mean.
- Without data access permissions and security controls, it is hard to scale into an enterprise AI platform.
In the end, the bottleneck in AI transformation is not model invocation cost or prompt engineering, but data quality, data definitions, and data governance.
If this is not resolved, AI may look smart, but in real work it becomes a tool that is hard to trust and delegate to.
2. KB Kookmin Bank Case: “Actionable AI” Comes from Trustworthy Data
KB Kookmin Bank pointed out the most realistic problems when introducing AI agents in finance.
Financial institutions have a lot of data, but that also makes it complex.
Customer data, transaction data, product data, consultation data, and document data are scattered across different systems.
On top of that, they must also consider the finance sector’s strong security, compliance, and internal controls.
The core of KB Kookmin Bank’s presentation was that “to move from searching AI to acting AI, you need more trustworthy data.”
An AI that only provides answers can tolerate some margin of error.
But an AI agent that actually performs tasks, supports decisions, and automates processes is different.
If AI acts based on incorrect data, financial and operational risks can arise immediately.
- Data silos must be reduced and data from multiple systems integrated.
- Structured and unstructured data must both be handled.
- Business metadata must clarify the meaning of data.
- Data that AI can access and data it must not access must be distinguished.
- Finance-level data governance must also be applied to AI agents.
This case shows what large enterprises and financial institutions should do first when introducing AI.
Rather than adding more AI models, they should first build a data foundation that AI can trust and act on.
3. Semantic Views: The Core Layer That Lets AI Understand Company Data
One of the most important concepts at this event was semantic views.
A semantic view is a layer that turns data from mere numbers and columns into information carrying business meaning.
For example, even the word “sales” can mean different things by department.
- For the sales team, sales may be based on contract date.
- For finance, sales may be based on accounting recognition date.
- For marketing, sales may mean campaign-attributed sales.
- For overseas business teams, the exchange-rate application 기준 may differ.
People can ask for context and adjust during meetings.
But if AI has no clear definitions, it may generate the most plausible answer.
That is the most dangerous point in enterprise AI.
Semantic views reduce this problem.
They define “which table and which column sales comes from, and under what aggregation rules,” in a way AI can understand.
It is a way of encoding business rules into data, such as “new products are products launched within three months,” “North America means the United States and Canada,” and “active customers are customers who purchased within the last 30 days.”
Ultimately, semantic views can be seen as the new enterprise dictionary for the AI era.
This dictionary must be built well for AI to answer business questions accurately.
4. Amorepacific Case: Building AI-Ready Data with Semantic Views and Ontology
Amorepacific is a representative retail company with brands, channels, countries, and product lines intricately intertwined.
In a company like this, even with lots of data, answering a single question is not easy.
For example, a question like “Why did North American new product sales increase last quarter?” requires multiple datasets.
- You need the definition of a new product.
- You need the scope of North America.
- You need sales criteria by channel.
- Promotion, inventory, and customer response data are also needed.
- You must understand the relationships among products, customers, inventory, and performance.
Amorepacific did not stop at simply gathering data in one place.
Through a semantic manager, it managed domain terms, semantic views, deployment history, and metadata.
It also added a search function based on Cortex Search so that differently written product names, such as Sulwhasoo, could still be mapped to the correct brand.
On top of that, it applied ontology.
Ontology is a technology for expressing relationships between data objects.
For example, when an inventory shortage occurs, it does not stop at simply saying “inventory is low.”
It can show, in a relationship-centric way, which product’s demand increased, which channel saw higher sales, which promotion had an effect, and whether a stockout may occur afterward.
The important point in this case is that it created an environment where business staff can ask questions without knowing code.
Data analysis has shifted from being the work of a few specialists to a working method that improves enterprise-wide productivity.
5. GS Construction Case: Reducing Data Preparation from Two Weeks to One or Two Days
The GS Construction case showed how AI can change field operations.
The construction industry is one where data structures differ by project and information accumulates at the site level.
Previously, staff had to enter the system, find menus, set conditions, download to Excel, and process it again.
For example, to answer “Which sites had welding work yesterday?” it required several steps.
But with Cortex Agents and semantic views, users can ask in natural language and get answers immediately.
- AI interprets the meaning of the question.
- It finds the necessary data based on the semantic view.
- It queries related data such as L2 data.
- Business staff review the result and make a decision.
GS Construction explained that in the process of building AI-ready data, tasks that previously took more than two weeks were reduced to a day or two.
This is not merely workflow automation.
It is a signal that even in an industry thought to be conservative, AI can directly improve enterprise productivity.
6. Coco (Cortex Code): A Natural-Language Coding Tool That Lowers the Barrier to Data Analysis
Another change emphasized by Snowflake was Coco.
Coco is short for Cortex Code, a coding assistant that helps with data tasks using natural language.
In the past, data engineers had to build SQL, Python, pipelines, and dashboards manually.
Now, the direction is shifting toward AI creating dashboards and analysis flows when you type something like, “I want to visualize this data like this.”
This matters because the user of data is changing.
Previously, the workflow was to request something from the data team and wait.
Going forward, business staff will directly ask questions, build dashboards themselves, and obtain insights themselves.
For companies, this change means more than just cost reduction.
When data analysis bottlenecks decrease, decision-making becomes faster.
Decision speed is directly linked to sales, cost, and risk management.
That is why AI data platforms are becoming infrastructure investments directly tied to enterprise competitiveness rather than simple IT investments.
7. WEMADE PLAY and Neowiz Cases: Making Game Data Understandable to AI
Game companies are an industry where massive amounts of log data flow in every second.
User behavior, payments, logins, stage progression, ad responses, and churn signals keep accumulating.
The problem is that this data often arrives in JSON format that is difficult for people to read.
WEMADE PLAY applied three steps to transform game logs into a structure AI can understand.
- First, it flattened complex log structures into column form.
- Second, it created a single source so scattered data could be answered from one place.
- Third, it applied semanticization, adding descriptions and meaning to the data.
Once this is done, game planners or operators can ask AI questions without knowing complex queries.
For example, when asked, “Where did the recent payment conversion rate drop?” AI can understand the meaning of the logs and explore the data.
Neowiz solved another problem.
Because multiple game IPs are operated independently, each team has different analysis requirements.
Previously, multiple game teams shared the same infrastructure, creating contention and operational burden.
Neowiz used Snowflake to unify governance while keeping analytics compute independent for each game.
The advantage of this structure is clear.
Security and data governance are managed under a single standard, while each game team can freely analyze according to its own workload.
It is a case that achieves both data control and analytical freedom.
8. Banksalad Case: For Small Data Teams, Automation Is a Survival Strategy
Banksalad receives a large amount of data every day due to the nature of its financial app.
But with not enough data engineers, they had to handle existing pipeline operations, new data ingestion, quality checks, and incident response all at once.
In particular, repetitive batch recovery work was a major burden.
In the previous setup, Airflow coordinated execution, PySpark jobs on EKS performed transformations and aggregations, and results were stored in S3.
To improve cost efficiency, they used spot instances, but when spot instances were reclaimed, batch jobs failed.
When a batch failed, staff had to inspect logs and recover it manually.
If this repeats, the data team spends time on incident recovery instead of building new data use cases.
Banksalad reduced the operational burden by building an automated environment based on Snowflake.
The key was creating a self-service environment where errors are detected immediately and the business team can verify data quality directly.
This case shows the direction of data engineering in the AI era.
If data engineers are tied up with infrastructure failures and repetitive recovery, AI adoption cannot scale.
Going forward, the core of data engineering is not writing more code, but building a more stable and automated data operations system.
9. Data Streams: A Real-Time Data Processing Approach That Reduces Kafka Operational Burden
Real-time data processing is becoming increasingly important in the AI era.
Factory sensor data, security events, user behavior logs, payment events, and app click data are difficult to handle only in batch.
Many companies have used Kafka to process this kind of data.
But Kafka is powerful and also difficult to operate.
Before creating topics, cluster sizing must be done, and brokers, partitions, replication, rebalancing, and connectors must be managed.
What the data team wants is real-time data and insights, but in practice much time is spent operating infrastructure.
Snowflake Data Streams is an approach to reduce this problem.
The starting point is not the cluster, but the topic.
Topics are managed as catalog assets inside Snowflake, while the physical data movement servers are separated into broker groups.
- It reduces the burden of complex cluster design.
- It handles real-time data around topics.
- It reduces connector configuration overhead.
- It allows focus on data freshness and pipeline status.
- It applies the concept of separating compute and storage to streaming as well.
This trend means real-time data analytics are no longer the domain of a few big tech firms, but part of the AI infrastructure of ordinary companies.
10. Nepes Case: In Manufacturing AI, a Structure of “Trust It, But Never Hand It Over Blindly” Is Important
Nepes is a company that manufactures semiconductor components and chemical products.
When using AI agents in manufacturing, greater caution is needed than in ordinary office work.
That is because equipment, quality, production schedules, defect rates, and safety issues are directly connected.
Nepes created a structure that uses structured and unstructured data together.
For structured data such as MES, ERP, and equipment data, it uses Cortex Analytics, while for unstructured data such as manuals and inspection records, it uses Cortex Search.
Especially important is the human-in-the-loop structure.
Even if AI analyzes and recommends, humans intervene at points where final judgment is needed.
In manufacturing, blindly handing things over to AI is not the answer.
The key is to clearly design what should be automated and where humans should verify.
This case shows the practical direction of industrial AI.
AI is likely to spread first as an assistive decision-making system that increases the speed and accuracy of on-site judgment, rather than fully replacing people.
11. Tossplace Case: An Apache Iceberg Strategy That Keeps Data Intact and Simply Adds the Right Engine
One of the most important technical keywords at this event was Apache Iceberg.
Apache Iceberg is an open table format that lets multiple engines read the same data without locking it into a specific platform.
In traditional data platforms, if system A’s data had to be used in system B, it had to be copied or moved.
That creates data duplication, higher storage costs, synchronization issues, and debates over consistency.
The problem becomes even bigger in AI agent workloads, where it is hard to predict when and how many queries will occur.
Tossplace solved this problem by “keeping one copy of the data and attaching the engine suited to the workload.”
Below, there is only one Iceberg table in object storage.
Above it, Polaris Catalog manages the latest snapshot, location, and permissions.
Snowflake handles batch ETL, BI, and ad hoc analysis, while StarRocks handles serving workloads that require high concurrency.
- Data replication pipelines are reduced.
- Since both engines read the same snapshot, consistency disputes decrease.
- You can choose the most suitable engine by workload.
- The risk of vendor lock-in can be reduced.
- Cloud costs and operational complexity can be managed simultaneously.
This is a core direction for next-generation data platforms.
Going forward, companies are more likely to manage data with open standards and combine multiple analytics engines rather than putting everything into one giant platform.
12. KT Case: No Matter the Tools, the Metrics Must Be One
KT addressed the problem of standardizing data definitions in a large organization.
KT’s BIDW is a massive data environment used by a large share of employees, with more than 200,000 reports active.
In such an organization, the same “sales” or “subscriber count” can have different calculation standards by department.
The problem is that meetings begin with aligning on what the numbers mean.
If one team includes certain items and another excludes them, the same metric can lead to different conclusions.
When AI enters this environment, the confusion can grow even more.
KT’s principle was “no matter the tools, the metrics and definitions must be one.”
Even if different environments such as Snowflake, Databricks, and existing analytics tools are used, they should reference the same semantic definitions.
Without this, even the semantic layer can become fragmented by tool.
The KT case shows the most important real-world issue in enterprise AI transformation.
To use AI well, the organization’s numerical language must first be unified.
If metric definitions are not unified, attaching AI agents can spread incorrect answers faster.
13. Conditions for an Enterprise AI Platform: Scalability, Openness, Governance
The conditions for an enterprise AI platform that can be summarized from this event fall into three main categories.
First, scalability.
It must work stably whether there are 10 users or 10,000.
Data ingestion, processing, BI, AI agents, and real-time serving workloads must not interfere with one another.
Once AI spreads across the enterprise, usage becomes hard to predict.
So compute resources must be separable by workload and flexibly scalable.
Second, openness.
It must support open standards such as Apache Iceberg.
If data is locked into a vendor-specific format, long-term cost and strategic flexibility decline.
In the future, companies are likely to use more than one tool and combine the best engine for each task.
Third, governance.
Data is a company’s core asset.
Who can see which data, which AI agents can read which data, and how sensitive information is masked must all be managed.
In the past, it was enough to manage human access permissions, but now AI agents must also be managed like users.
In particular, sensitive data such as HR, compensation, customer information, financial information, and trade secrets must be more tightly controlled during AI training and response generation.
If this is not organized, companies may want to expand AI but end up stopping because of internal risks.
14. The Core Point Rarely Stated in Other News: AI ROI Depends More on ‘Data Operating Cost’ Than on Model Cost
One of the most important but relatively underreported points at this event is the structure of AI ROI.
Many news stories focus on AI model competition, GPU investment, and big-tech cloud competition.
But in real enterprise settings, the key factor determining AI return on investment is a little different.
For a company to make money with AI, it must reduce three types of cost.
- First, it must reduce the time people spend finding and organizing data.
- Second, it must reduce cloud costs for duplicate storage and data movement.
- Third, it must reduce decision-making costs caused by metric inconsistencies and data quality issues.
AI model invocation cost is easy to see.
But the truly large costs are the hidden costs that arise when data is not prepared.
The cost of aligning numbers in every meeting, requesting data from the data team, recovering pipeline failures, and copying the same data into multiple places keeps accumulating.
That is why semantic views, Iceberg, data streams, and unified governance matter.
These technologies are not for looking impressive; they are technologies that lower operating costs in the AI era and improve enterprise productivity.
In the end, companies that use AI well are not the ones that built a better chatbot first.
They are the ones that first fixed data definitions, data quality, data access permissions, and data movement structures.
15. AI Data Platform Trends from an Economic Perspective
This event also has significant meaning from the perspective of the global economy and industrial structure changes.
Companies adopt generative AI not because it is a passing tech trend.
It is because AI is the most realistic means of increasing enterprise productivity under high interest rates, rising labor costs, and low-growth pressure.
But adopting AI does not automatically lead to productivity gains.
Companies whose data is not organized get only limited automation benefits even if they add AI.
Conversely, companies with a well-organized data foundation can achieve results much faster even with the same AI model.
This difference is likely to be reflected in company value and investment decisions going forward.
Companies with strong AI infrastructure can gain an advantage in cost efficiency, decision speed, and customer responsiveness.
By contrast, companies with siloed data may see AI investment costs rise while results lag.
So the key metric for digital transformation going forward is not simply “Did we adopt AI?”
It is more likely that the more important question will be, “Do we have a data structure that AI can understand?”
16. A Checklist Companies Should Review Right Now
If we translate the event into practical company terms, it can be summarized with the following questions.
- Are our definitions of sales, customers, products, inventory, and performance the same across departments?
- Are data that AI can access and data it must not access clearly separated?
- Can structured and unstructured data be searched and analyzed together?
- Can business staff ask questions in natural language without help from the data team?
- Can data be managed as a single source without copying it across multiple platforms?
- Is there a structure that reduces the operational burden of real-time data pipelines?
- Have we verified AI agents’ actions and defined where humans should intervene?
- Even if multiple analytics tools are used, are metric definitions kept as one?
If you cannot answer these questions, your AI project may struggle to deliver results.
Conversely, companies that can answer them are much closer to being ready to attach AI agents to real work.
< Summary >
The core takeaway of Snowflake World Tour Seoul was that “AI performance depends more on data readiness than on the model.”
KB Kookmin Bank emphasized that a trustworthy data foundation is needed for actionable AI.
Amorepacific built a data structure that AI can understand through semantic views and ontology.
GS Construction reduced AI-ready data preparation time from two weeks to one or two days.
WEMADE PLAY and Neowiz transformed game logs and large-scale analytics environments to be more AI-friendly.
Banksalad reduced the repetitive operational burden on a small data team through automation.
Nepes showed a human-in-the-loop structure in manufacturing where AI is trusted but never handed over blindly.
Tossplace created a structure using Apache Iceberg where data remains in one copy and engines are attached as needed.
KT showed the core of large-scale data governance: no matter the tools, metric definitions must be one.
In conclusion, the companies that use AI well are the ones that first fixed data definitions, quality, governance, and open data structures.
[Related Articles…]
- The Core Strategy Transforming Enterprise Productivity in the AI Agent Era
- How Data Infrastructure Innovation Is Powering the Next Wave of Digital Transformation
*Source: [ 티타임즈TV ]
– AI 잘 쓰는 회사들이 먼저 고친 것 (스노우플레이크 월드투어 서울 총정리)


