2026 marks a critical turning point for the AI industry.
As large models move from technical demonstrations to production systems, and Agents advance from proof-of-concept to large-scale deployment, a consensus is forming: the decisive factor in AI competition has shifted from models and computing power as general capabilities to the data infrastructure that supports AI operations.
Recent McKinsey research shows that although the vast majority of enterprises have deployed AI projects, fewer than 1% have truly achieved large-scale value realization. The core challenge no longer lies in general capabilities such as models and computing power, but in systemic issues including high-quality data supply, industry scenario implementation, and industrial collaboration.
Data is becoming the "last mile" of AI industrialization.
Value—data silos and quality issues are the core obstacles. A survey by the China Academy of Information and Communications Technology covering more than a hundred enterprises shows that the primary challenge for intelligent implementation is insufficient data quality (48.7%), followed by insufficient data governance capability (45.4%). More than 70% of AI model training time is spent on data cleaning and preprocessing.
Fei-Fei Li recently extended this judgment from language models to embodied intelligence: "Robots are extremely data-poor, and both training and evaluation data are severely insufficient. Large language models only need to understand semantics, but robots must understand the physical world—weight, friction, deformation, as well as spatial relationships and obstacle layouts." In fact, as early as 2006 she said: the bottleneck of computer vision is not in models, but in data.
Twenty years later, this statement has been repeatedly validated at every stage of AI industrialization.
Policy Resonance: From "AI+" to "Model-Data Resonance"
In China, a policy-level "data awakening" is accelerating.
In April 2026, the Ministry of Industry and Information Technology and the National Data Administration jointly issued the Notice on Jointly Implementing the 2026 "Model-Data Resonance" Action, proposing seven key tasks including building industry general-knowledge datasets, creating distinctive intelligent agents, and establishing "Model-Data Resonance" spaces, with the goal of basically forming a virtuous cycle of "data-model-scenario application" by the end of 2026.
This policy evolution from "AI+" to "Model-Data Synergy" marks a shift in the focus of large-scale AI implementation from the model side to the data side. The policy explicitly requires: sorting out data resources by industry and refining industry general-knowledge high-quality datasets, no fewer than 5 per industry; building industry-specific datasets for high-value scenarios and creating specialized models or distinctive intelligent agents; establishing and improving evaluation datasets to form a virtuous cycle of "evaluation diagnosis—targeted dataset optimization—model capability improvement."
Meanwhile, national data infrastructure construction is accelerating. The National Data Administration has announced 24 trusted data space pilots, with the goal of completing 30 pilot constructions in 2026. Data elements are moving from "resources" to "assets," and from "accumulation" to "circulation." On the market side, according to IT market consulting agency forecasts: China's AI data infrastructure market was approximately 45 billion yuan in 2025 and is expected to reach 198.4 billion yuan in 2029, with a compound growth rate of approximately 44.9%. 2026 will enter a concentrated construction period, with a scale of 72.9 billion yuan and a year-on-year growth rate of 62.0%. The dual drive of policy and market is creating a historic window of opportunity for the "China solution" of AI data infrastructure.
KeenData: A "China Sample" of Next-Generation AI Data Infrastructure
KeenData anchors itself in the AI data infrastructure track and pioneered the "AI-in-Lakehouse intelligent driving architecture." This architecture natively integrates the lakehouse engine, multimodal computing engine, and training-inference acceleration engine, enabling unified access, intelligent governance, annotation, and operation of structured and unstructured data, making data directly consumable by AI.
In terms of product matrix, KeenData has built a clearly layered full-stack product system. At the bottom is the KeenData Lakehouse multimodal lakehouse platform, providing foundational capabilities such as data governance, multimodal computing, and AI training; at the top is the KeenAgenticOS intelligent agent development operating system. Together they form the KeenData Agentic Lakehouse Platform intelligent foundation, opening up a complete closed loop of "data-action-operations"—through KeenClaw, intelligent agents leap from single-point tools to organization-level productivity; through KeenRouter, the transition from simple model invocation to AI consumptive operations is achieved, with support for precise Token billing. As a result, KeenData has formed complete full-stack coverage of AI data infrastructure.
Intelligent Agents Across Thousands of Industries
From the perspective of implementation practice, KeenData has served more than 300 large organizations, covering core industries such as energy, finance, manufacturing, and government. A large energy state-owned enterprise, through the AI data foundation built by KeenData, integrated 61 core systems and 1.2PB of data, reducing the efficiency of viewing business analysis reports from 1 week to 4 hours. In the financial sector, China CITIC Bank jointly built a full-domain real-time data foundation with KeenData, and the AI-driven credit approval system compressed processing time from days to minutes. In the government sector, KeenData has supported the access of more than 1,000 data entities and the release of more than 2,000 data products.

The Sinopec data foundation was listed by the State-owned Assets Supervision and Administration Commission as a benchmark for central enterprise digital transformation, and the FAW Hongqi digital intelligence foundation construction plan was selected as a best practice case for central state-owned enterprise digitalization.
From the perspective of global layout, KeenData's overseas expansion is accelerating. It has established customer networks and business cooperation in multiple countries including Saudi Arabia, Singapore, Japan, Malaysia, the Philippines, and South Africa, and has set up an overseas company in Saudi Arabia.
Next Stop: From "Data Management" to "Intelligent Productivity"
"Without a solid data infrastructure, there is no true intelligent economy." KeenData Chairman Yu Yang has repeatedly emphasized this judgment.
As AI moves from laboratories to workshops, offices, and every corner of urban governance, what truly determines success or failure is whether the AI data foundation is solid. From "lakehouse" to "AI-in-Lakehouse," from "data management" to "intelligent productivity," the reconstruction of a new generation of AI data infrastructure is underway.
Whoever can truly organize data, knowledge, and business will be able to navigate the cycle in the deep water zone of the AI industry and reach the true shore of the intelligent economy.
News & Updates

As AI Moves Toward Large-Scale Application, the Long-Term Value of Data Infrastructure Is Emerging

2026 Tomorrow.City · Shanghai Opens | KeenData Supports Global Urban Intelligent Evolution with AI Data Infrastructure

AI Global Competition Enters Data Depth | KeenData Invited to Singapore Industrial Cooperation Exchange Meeting
