AI & Technology

AI’s pivotal moment in multimedia is now

By Hamed R. Tavakoli, Head of Visual AI Systems Research, Nokia

Much of the conversation around artificial intelligence (AI) has focused on its impact on productivity and automation. Less attention has been paid to how it will transform the way [H(1] content is created, delivered and consumed – despite its potential. As AI adoption scales and places new demands on digital infrastructure, multimedia is becoming an increasingly important part of that broader shift.

Imagine watching a film at home where the picture adjusts scene by scene to preserve detail and lighting. Or imagine watching a live football match that brings forward the moments, angles or players you care about most as the game unfolds.

These are not distant ideas, but early signs of a broader shift taking place across multimedia, with AI at its centre. Across creation, editing, management and analysis, AI is enabling more personalised and interactive experiences, as well as increasingly rich immersive environments.

Realising this potential at scale, however, is not straightforward. Unlocking the full benefits of AI in multimedia requires greater standardisation. Today, the integration of complex machine learning (ML) models varies across platforms, hardware and implementations, creating a fragmented landscape. Without a more consistent approach, there is a risk that AI-driven tools will struggle to work together seamlessly.

Enabling interoperable AI innovation across multimedia  

Thousands of innovations and iterations have been embedded across multimedia systems over recent years, each driving the industry forward one step at a time.

Video streaming platforms, for instance, are experimenting with AI-driven compression, while creation and editing tools are using generative models to transform user ideas into rich media. Across audio and visual services, more personalised and adaptive experiences are emerging.

As AI becomes more widely embedded across multimedia, ensuring these technologies work seamlessly together will be increasingly important. Innovation is moving quickly across platforms and ecosystems, creating new opportunities for interoperability and collaboration.

Two video platforms may use different AI-based compression methods, meaning content requires additional processing to optimise playback across services and devices. Common approaches could help reduce these inefficiencies, supporting smoother interoperability and more predictable performance across multimedia systems.

This is where initiatives such as MPEG-AI play an important role, helping establish shared technical foundations that support compatibility across systems while enabling continued innovation across the industry. In this sense, standards are not only a technical enabler, but also part of the wider infrastructure needed to support the next wave of AI innovation.

The vision for MPEG-AI bridging AI and multimedia

Effectively serving as an umbrella for standardisation, MPEG-AI is being developed as a framework to bring greater consistency to how AI and multimedia interact. At its core, MPEG-AI aims to consider two complementary aspects: how AI can be used to enhance multimedia systems, and how multimedia can be structured to better serve AI-driven processes.

These two aspects are often described as “AI as a multimedia coding tool”, and “multimedia for consumption by AI systems”. Together, they provide a foundation for exploring how AI can be more effectively integrated across the multimedia industry.

The vision for MPEG-AI comes to life through its application. First, take AI as a powerful multimedia coding tool. Through neural networks, AI can enhance multimedia coding by applying deep learning models to compress and improve the quality of video. MPEG-AI aims to standardise the use of AI-driven compression methods, helping enable interoperability across platforms.

At the same time, multimedia data – such as videos or images – can be structured in ways that make it more accessible to AI systems, supporting tasks like object recognition or object tracking. MPEG-AI is expected to play a key role in defining how multimedia is encoded and structured to be more AI-friendly. In facilitating feature coding for machine consumption, it has the potential to transform an input feature map, obtained from a neural network, into a decodable bitstream for any machine task.

Looking further ahead, the MPEG-AI family of standards will continue to offer improved digital experiences for consumers over the coming years. People will benefit from AI enhanced applications across various devices, such as more immersive gaming experiences with virtual reality, mixed reality, or a supercharged metaverse.

Outside of the home, AI in multimedia could strengthen smart cities and improve agentic AI use cases.

Innovations that have long been talked about could become reality in months rather than years – rolled out to the masses thanks to the umbrella family of standards, MPEG-AI.

What’s next for MPEG-AI?

While much of today’s focus is on AI-enhanced content experiences, attention is also turning to the infrastructure required to support AI-driven multimedia systems at scale. One growing area of focus is tensor data, including the intermediate feature data generated within neural networks.

This has important implications for emerging applications such as split inferencing, where AI workloads are shared between edge devices and cloud infrastructure. In these scenarios, part of a neural network can run locally on a device, with compressed feature data transmitted to the cloud for further processing. Improving how this data is handled and delivered could help enable more scalable and efficient AI-powered multimedia services.

Progress is already being made in this area. Recent developments in neural network coding and feature coding for machines are exploring how tensor compression standards can support a broader range of use cases beyond model parameters alone. This includes applications such as machine-driven feature coding, federated learning and emerging immersive media formats including Gaussian splats and 3D scene representations.

Alongside efficiency and scalability, future standards development is also expected to support greater reliability and trust in AI systems, including mechanisms for verifying the authenticity of neural network data and updates.

The industry now sits at a key inflection point for AI as a multimedia coding tool and multimedia for consumption by AI. Establishing common frameworks, such as MPEG-AI, will be an important step in supporting a more cohesive and scalable future for AI-driven multimedia.

Author

Related Articles

Back to top button