Enfabrica Corp.’s hybrid memory fabric system designed to improve efficiencies in large-scale distributed, memory-bound AI inference workloads is now available.
Called EMFASYS, the hardware/software platform is claimed to be the first commercially available system that integrates high-performance remote direct memory access (RDMA) Ethernet networking with parallel ComputeExpressLink (CXL) based DDR5 memory channels.
Enfabrica said the solution provides AI compute racks with elastic memory bandwidth and memory capacity.
The EMFASYS platform provides higher utilization of GPU and high-bandwidth memory resources in the compute rack while scaling to high user/agent count, accumulated context and token volumes. As AI workloads are growing substantially — requiring 10 to 100 times more compute per query than the previous large language model (LLM) — there are billions of batched inference calls per day across AI clouds, the company said.
“AI Inference has a memory bandwidth-scaling problem and a memory margin-stacking problem,” said Rochan Sankar, CEO of Enfabrica. “As inference gets more agentic versus conversational, more retentive versus forgetful, the current ways of scaling memory access won’t hold.”
Other features of EMFASYS platform include:
- 3.2 Tbps accelerated compute fabric SuperNIC
- Connects to up to 144 CXL memory lanes to 400/800 GbE ports
- Offloads GPU and HBM consumption
- Aggregates CXL memory bandwidth
- Read access times in microseconds
- Software-enabled caching hierarchy
- Lowers cost of LLM inference
- Outperforms flash-based inference storage
