ACE is a computing power engine designed for large-scale deep learning and intelligent computing.
Based on self-developed computing power card management technology, it improves the utilization rate of computing power cards, provides rich monitoring and operation methods, and various task scheduling strategies to help users establish computing power resource pools.
It provides comprehensive computing power management capabilities for AI model training, inference, simulation, rendering, bioinformatics analysis, numerical computing, and other scenarios.
ACE supports multiple heterogeneous GPU resources, allowing developers to flexibly choose combinations of card resources and CPU types according to their needs, in order to achieve optimal cost-effectiveness, compatibility with information and innovation, and other goals.
.png)
Heterogeneous GPU pooling
Satisfy the unified management of GPUs from different brands, virtualize and pool scheduling, support single card segmentation, single machine multi card, and multi machine multi card operation modes for training and inference scenarios, improve the efficiency of computing resource utilization, and cover mainstream GPU brands such as Huawei Ascend, Haiguang DCU, Tiantian Zhixin, Cambrian, and Denglin.
.png)
GPU resource monitoring
For expensive GPU resources, there is a need for rich monitoring data to support their full and efficient use, as well as fast fault perception and repair. ACE can provide fine-grained monitoring indicators for GPU/NPU/GPGPU and other computing card temperature, power, usage, video memory usage (amount), fan speed, card task count, card power limit mean, as well as historical review for various scale intelligent computing, supercomputing, and general computing hybrid scenarios. It can also output resource usage logs to interface with computing billing engines, supporting fine-grained management.

Automated high concurrency task scheduling
Support multi person batch concurrent use of tasks in training and inference scenarios, based on an enhanced real-time task scheduler that supports centralized, tiled, and nearby tasks FIFO、Gang、DRF、binpack、 Advanced scheduling strategies such as real load can meet the efficient management and allocation of resources by users in different scenarios.
.png)
Remote calling and network acceleration
Support calling the computing power resources of remote GPU servers through virtual GPU cards on hosts without GPU resources, achieving balanced utilization of CPU and GPU computing power resources.
Support IB/RoCE/RDMA high-speed network protocols, optimize network transmission efficiency and stability in high bandwidth and low latency scenarios, improve computing power cluster reliability, support DPU intelligent network cards, and optimize application offloading for cloud native scenarios.
Launch enterprise level intelligent agents to truly drive
business growth through AI
Professional consultants provide one-on-one services and tailor enterprise intelligent agent solutions