Imagination releases first E-Series GPU benchmarks showing AI upscaling and LLM prefill speed gains
Imagination's new E‑Series GPU demonstrates a 2.3 ms AI upscaling to double image resolution and a 4.7× faster LLM pre‑fill, according to LeiPhone and Jon Peddie Research.

Imagination has unveiled the first publicly available benchmark data for its upcoming E‑Series graphics processor. The results, supplied by independent test houses LeiPhone and Jon Peddie Research, focus on two headline capabilities: Neural Super Resolution (NSR) for image upscaling and the pre‑fill phase of large language models (LLMs) running on the same silicon.
The NSR technology, branded by Imagination as Neural Super Resolution, claims to double the resolution of a source image in just 2.3 milliseconds. The measurement was taken under controlled conditions using the LeiPhone and Jon Peddie Research test suites, which applied a standardized 4K‑to‑8K upscaling workload. No other timing data or comparative figures were disclosed beyond the 2.3 ms figure.
In parallel, the benchmarks show that the pre‑fill step for large language models—a memory‑intensive operation that loads model weights into GPU memory before inference—executes 4.7 times faster on the E‑Series than on Imagination's previous generation D‑Series. Both LeiPhone and Jon Peddie Research observed the same speed‑up factor across a range of model sizes, though the exact models and batch configurations were not enumerated.
Matrix‑Accelerator Architecture
A core element of the performance uplift is the inclusion of a dedicated matrix accelerator within the E‑Series GPU. The accelerator supports a suite of low‑precision data formats, including MXFP8, MXFP4, BF16 and FP4. These formats are designed to reduce the bit‑width of tensor operations, thereby increasing the throughput of matrix multiplications and convolution kernels that dominate AI workloads.
Imagination states that the matrix accelerator is tightly coupled to the shading engine. By embedding the matrix unit directly into the tile‑based deferred renderer (TBDR) pipeline, data movement between the graphics rasterizer and the AI compute blocks is minimized. LeiPhone’s analysis attributes up to a 39 % improvement in energy efficiency to this integration, though the exact measurement methodology is not disclosed.
Unified Software Stack
Another claim from Imagination is that the E‑Series can be programmed through a single software stack that spans traditional graphics APIs (OpenCL, Vulkan) and AI frameworks (Llama.cpp, ONNX, PyTorch). The company argues that this reduces the number of required drivers and shortens qualification cycles for system‑on‑chip (SoC) designers who need to support both rendering and AI inference on the same die.
The unified stack is meant to simplify development for embedded platforms where memory, power and silicon area are at a premium. By reusing the same driver infrastructure for both rendering and AI, OEMs could potentially avoid maintaining separate validation suites for graphics and neural‑network accelerators.
Bandwidth and Compression Benefits
Imagination also highlights the compression characteristics of the NSR model. According to the company, the NSR network is compressed by roughly 65 % of its original size. This compression translates into a 50 % reduction in memory‑bandwidth demand when compared with a native rendering path that does not employ NSR.
The reduced bandwidth requirement is especially relevant for SoCs that target high‑resolution displays but are constrained by memory‑interface limits. Imagination lists a maximum memory bandwidth of 768 GB/s for the E‑Series platform, but notes that actual performance will depend on how much of that bandwidth is allocated to graphics versus AI workloads.
While the compression lowers bandwidth pressure, Imagination acknowledges that the visual quality of NSR‑upscaled frames remains “close to” the reference native rendering output. No quantitative quality metrics such as PSNR or SSIM were provided in the benchmark release.
Limitations and Heterogeneous Architecture
Imagination is transparent about the fact that the E‑Series GPU does not replace dedicated neural‑processing units (NPUs) for all AI tasks. Workloads that require ultra‑low latency or highly specialized operators may still be better served by purpose‑built NPUs. The company therefore positions the E‑Series as one component of a broader heterogeneous compute fabric that includes CPUs, NPUs and other accelerators.
The benchmark documentation also stresses that real‑world impact will hinge on software‑engineer adoption and SoC memory configurations. The 4.7× LLM pre‑fill speed‑up, for example, assumes a memory subsystem capable of delivering the advertised 768 GB/s bandwidth. Systems with lower bandwidth may see reduced gains.
Similarly, the 2.3 ms NSR upscaling latency was measured under ideal test conditions. The documents caution that industrial deployment could encounter variability due to driver maturity, thermal throttling, or integration with other system components.
- Neural Super Resolution doubles resolution in 2.3 ms
- LLM pre‑fill 4.7× faster than D‑Series
- Matrix accelerator supports MXFP8, MXFP4, BF16, FP4
- Unified stack reduces driver count across OpenCL, Vulkan, PyTorch
The list above captures the four headline figures that Imagination and the independent test houses have published. All other performance observations remain qualitative or are expressed as relative improvements without absolute numbers.
From a practical standpoint, organisations evaluating the E‑Series for embedded or edge devices must weigh the documented speed‑ups against the uncertainties highlighted in the benchmark release. The energy‑efficiency gains tied to tile‑based rendering and matrix‑unit integration suggest lower power draw per frame, which could be valuable for battery‑operated products.
However, the dependence on a unified software stack means that development teams will need to verify compatibility with their existing toolchains. If a company already uses a dedicated NPU for latency‑critical inference, the E‑Series may complement rather than replace that hardware, requiring careful partitioning of workloads.
Memory bandwidth remains a decisive factor. Systems that cannot provision the full 768 GB/s may not realize the full 39 % energy‑efficiency improvement or the 4.7× LLM pre‑fill acceleration. Designers should therefore assess their memory subsystem early in the SoC planning phase.
Finally, the benchmark data does not guarantee broad industry adoption. Software vendors must integrate the new NSR model and matrix‑accelerator APIs into their rendering engines and AI frameworks before end‑users can benefit. Until that integration occurs, the performance numbers remain best‑case scenarios observed in controlled lab environments.
Sources
- Imagination E系列性能首秀:用AI把分辨率翻倍仅需2.3毫秒,Prefill性能提升4.7倍 | 雷峰网LeiPhone · September 28, 2026
- Imagination’s E-Series unifies graphics and AI – Jon Peddie ResearchJon Peddie Research · September 21, 2026


