Skip to content

Latest commit

 

History

History
159 lines (110 loc) · 13.6 KB

File metadata and controls

159 lines (110 loc) · 13.6 KB

Rec SDK

Ascend License Zread DeepWiki

✨ What's New

🔹 [2026.07.31]: Rec SDK 26.1.0 Release
🔹 [2026.04.25]: Rec SDK 26.0.0 Release

ℹ️ Introduction

Rec SDK offers an SDK for search, recommendation, and advertising services in the Internet market. It offers a framework for these services based on the Ascend platform to meet related model training requirements, thus supporting large-scale search-recommendation-advertising scenarios and facilitating efficient training of models for such scenarios.

Rec SDK provides the following features:

  1. Basic model training functions: Rec SDK supports single-node, single-card training and multi-node, multi-card distributed training.
  2. Recommendation-specific functions: Based on the sparse table solution, Rec SDK provides essential functions such as feature saving and loading, feature admission, and feature eviction.
  3. Large-scale sparse table functions: Rec SDK supports multi-level storage across accelerator card memory, host memory, and host drive. It also supports multi-node storage and dynamic scaling, with capacities exceeding 10 TB.

⚙️ Features

Component Feature Summary Documentation
tf_rec_v1 Supports single-/multi-node and single-/multi-card distributed training; provides recommendation-specific features such as feature saving and loading, feature admission and eviction, dynamic capacity expansion, dynamic shape, automatic graph rewriting, Hot_Embedding, customized WarmStart, incremental model saving/loading, single-table multi-query, and PCIe through; supports multi-level storage across accelerator memory, host memory, and host disk (capacity > 10 TB); provides performance and accuracy detection tools. Details
tf_rec_v2 Supports single-/multi-node and single-/multi-card distributed training; provides recommendation-specific features such as sparse table creation, query, saving and loading, and feature admission and eviction; supports large-scale sparse table storage. Details
torch_rec_v1 Supports single-/multi-node and single-/multi-card distributed training; provides recommendation-specific features such as hash mapping, EBC table lookup, Row-wise table sharding, pipelined lookup, and fused lookup operators; supports Row-wise distributed sparse table sharding. Details
torch_rec_v2 Supports single-/multi-node and single-/multi-card distributed training; provides recommendation-specific features such as hash mapping, Row-wise table sharding, dynamic sparse table expansion and eviction, and dynamic sparse table operators; implements dynamic sparse table operators based on the HKV high-performance key-value storage acceleration library. Details

🚀 Quick Start

Component Base Framework Adaptation Status Framework Type Description Documentation
tf_rec_v1 TensorFlow Non-fully-offloaded Sparse recommendation framework Non-fully-offloaded sparse recommendation framework based on TensorFlow, adapted for NPU devices Quick Start
tf_rec_v2 TensorFlow Fully-offloaded Sparse recommendation framework Fully-offloaded sparse recommendation framework based on TensorFlow, adapted for NPU devices (PoC) Quick Start
torch_rec_v1 PyTorch + TorchRec Non-fully-offloaded Sparse recommendation framework Non-fully-offloaded sparse recommendation framework based on open-source PyTorch and TorchRec, adapted for NPU devices Quick Start
torch_rec_v2 PyTorch + TorchRec Fully-offloaded Sparse recommendation framework Fully-offloaded sparse recommendation framework based on open-source PyTorch and TorchRec, adapted for NPU devices (PoC) Quick Start

Key terminology:

  • Non-fully-offloaded: A hybrid mode in which some computing tasks are executed on the NPU and others on the CPU.
  • Fully-offloaded: A mode where all computing tasks are offloaded to the NPU for execution.
  • PoC: Proof of Concept, indicating that the component is still in the experimental verification phase and features may be incomplete or unstable.

📦 Installation Guide

The following product models are supported by Rec SDK:

  • Atlas A2 Train Series
  • Atlas A3 Train Series
  • Ascend 950PR&950DT Series
Component Installation Guide
tf_rec_v1 Installation Guide
tf_rec_v2 Installation Guide
torch_rec_v1 Installation Guide
torch_rec_v2 Installation Guide

📘 Usage Guide

Rec SDK provides comprehensive usage and development documentation to help you understand the architecture, tuning, and usage of each component. For details, refer to the Ascend community recommendation development documentation.

Mode adaptation examples:

Model Framework Component Code link
DIN PyTorch torch_rec_v1 Code link
DLRM(DCNv2) PyTorch torch_rec_v1 Code link
GR PyTorch torch_rec_v1 Code link
GR PyTorch torch_rec_v1 Code link
MMOE, ETA PyTorch torch_rec_v1 Code link
GR PyTorch torch_rec_v2 Code link

🗺️ Roadmap

Roadmap (2026Q3)

Roadmap (2026Q2)

Roadmap (2026Q1)

🔀 Version Maintenance Strategy

The maintenance stages of Rec SDK version branches are as follows:

Status Duration Description
Plan 1–3 months Planned features
Development 6–12 months Develop new features and fix issues, and release new versions regularly. Different strategies apply to different PyTorch versions: regular branches have a 6-month development cycle, and long-term support branches have a 12-month development cycle.
Maintenance 1 year / 3.5 years Regular branches are maintained for 1 year, and long-term support branches for 3.5 years. Major bugs are fixed, no new features are merged, and patch versions are released based on bug impact.
End of Life (EOL) N/A The branch no longer accepts any changes

torch_rec_v1 Version Maintenance

The maintenance status of each torch_rec_v1 version is as follows:

torch_rec_v1 version Maintenance policy Current status Release date Next status EOL date
1.1.0 Regular branch Maintenance July 23, 2025 Expected to enter unmaintained status after September 30, 2026 September 30, 2026
1.2.0 Regular branch Maintenance January 4, 2026 Under maintenance

🛠️ How to Contribute

We welcome your contributions. For the contribution process and specifications, see the Contribution Guidelines. Before contributing, sign the Open Project Contributor License Agreement (CLA).

  1. If you encounter a bug, submit an issue.
  2. If you plan to contribute bug fixes, submit a pull request (PR). See Contribution Requirements.
  3. If you plan to contribute new features or functionality, create an issue to discuss it with us first. Describe the background or purpose of the requirement, the design, and its impact on existing APIs. Submitting a PR without prior discussion may lead to rejection, as the evolution direction of the project might differ from your ideas.

⚖️ Related Information

🔹 Release notes
🔹 License
🔹 Document license
🔹 Disclaimer
🔹 Component-related information

Component FAQ Security hardening
tf_rec_v1 FAQ Security hardening
tf_rec_v2 FAQ Security hardening
torch_rec_v1 / Security hardening
torch_rec_v2 / Security hardening

🤝 Suggestions and Communication

You are welcome to raise questions and join discussions through the following channels.

Resource Description
Create an issue Submit a bug, requirement, or suggestion
Community tasks View and claim community tasks
Meeting calendar Regular community meetings and events

🙏 Acknowledgments

Rec SDK is jointly contributed by the following Huawei departments:

  • Ascend Computing Application Enablement Development Dept
  • Software Platform Dept, Computing Product Line
  • UnifiedBus Computing Cluster Development Dept
  • Technology Development Dept, Computing Product Line
  • Poisson Lab

Thank you to everyone in the community for your PRs. We warmly welcome your contributions to Rec SDK!