The world of artificial intelligence is seeing a significant shift. While large language models, or LLMs, like the technology behind ChatGPT, have captured public attention with their ability to generate text and code, new research indicates a strong push toward "agentic AI." This isn't just about AI understanding language, but about it actively interacting with and changing the external world, moving from a conversational partner to a system that can take action.

A key development in this evolution involves new model architectures. Traditional LLMs are "autoregressive," meaning they generate text one word or token at a time, from left to right. However, a new class of technology, Diffusion Language Models, or DLMs, offers a different approach. DLMs refine tokens through an iterative denoising process, allowing them to update multiple uncertain tokens in parallel and use context from both directions. This flexibility is particularly appealing for "edge agents," which are AI systems deployed on devices like smartphones or smart sensors, rather than in large data centers. DLMs can potentially reduce response delays and communication overhead, making AI more robust and efficient in real-world, dynamic environments.

This move to agentic AI means that language models are no longer just producing outputs for humans to read. Instead, they are being integrated into systems that allow those outputs to directly alter external states. This includes AI models that can call tools, operate digital interfaces, delegate tasks, and even control robots or laboratory equipment. This represents a significant leap from simply answering questions to actively performing tasks, whether in digital, social, virtual, or physical environments.

However, the path to fully autonomous, real-world AI agents is still long. While there is clear progress in expanding the interfaces through which AI can act, the evidence for robust task completion, error recovery, proper authorization, or independent verification is less convincing. Current improvements in interoperability, like the Model Context Protocol, facilitate communication between AI systems but do not inherently establish trustworthy delegation. Even multi-agent organizations, where different AIs specialize in tasks, can add complexity and correlated failures.

One area where agentic AI is being explored is in managing complex infrastructure, such as heating, ventilation, and air conditioning, or HVAC, systems in buildings. These systems generate vast amounts of sensor data, but extracting useful insights can be challenging due to inconsistent naming conventions and fragmented documentation. LLMs are being investigated to help normalize data, provide operator support, assist with building energy modeling, and offer advisory interfaces for physics-based controllers. This could lead to more efficient building operations and significant energy savings.

Despite the promise, the deployment readiness for LLM-based HVAC operations is still nascent. A review of 66 peer-reviewed studies published between 2023 and March 2026 found that most research is still conceptual or experimental. Only a handful of studies reached pilot-level evidence, and none reported sustained operational deployment. While some specific, human-in-the-loop applications, such as point-name normalization and document-grounded operator support, are considered near-term possibilities for trials, widespread industry adoption is not yet on the horizon.

Project Ares analysis suggests that the true impact of agentic AI will hinge on bridging the gap between theoretical capabilities and practical, reliable deployment. The development of DLMs addresses some fundamental limitations of traditional LLMs for edge devices, potentially making AI agents more ubiquitous and responsive. However, the critical challenge lies in building robust "harnesses" around these models – the surrounding systems that manage their actions, ensure safety, and provide mechanisms for human oversight and intervention. Without these safeguards, the promise of agentic AI could be overshadowed by concerns about uncontrolled or unpredictable actions, particularly in sensitive physical environments.

What to watch next is how these different threads converge. We will need to see more real-world pilot programs, particularly in areas like building automation and robotics, that move beyond theoretical models and demonstrate sustained, safe, and effective operation. The focus will be on the development of reliable integration platforms and robust safety protocols that allow AI to act with increasing autonomy without sacrificing human control or accountability. The evolution from language models to world-acting systems is underway, but the journey from research to widespread, trustworthy deployment is just beginning.