AFD-Ledger: Analytical Provisioning for Attention-FFN Disaggregation
AFD-Ledger is an analytical provisioning system that operates offline, aimed at determining if Attention-Feed-Forward Network (FFN) Disaggregation (AFD) achieves greater throughput compared to collocated setups for Mixture-of-Experts (MoE) language models. This system autonomously provisions AFD and collocated setups by utilizing an analytical execution model alongside a hardware search that is evaluation-bounded. By optimizing hardware allocation and deployment structure together, it circumvents the high costs associated with exhaustive provisioning. The paper, found on arXiv (2608.04502), explores a previously unanswered question regarding AFD systems: does AFD exceed the performance of the best collocated deployment under identical conditions? AFD-Ledger's analytical framework allows for effective comparisons in deployment scenarios where exhaustive provisioning is practical, making it significant for the AI technology field focused on enhancing large language model services.
Key facts
- AFD-Ledger is an offline analytical provisioning system.
- It compares AFD (Attention-FFN Disaggregation) with collocated deployments.
- Targets Mixture-of-Experts (MoE) language models.
- Uses an analytical execution model and evaluation-bounded hardware search.
- Jointly optimizes hardware assignment and deployment organization.
- Addresses a deployment question: does AFD provide higher throughput under same conditions?
- Paper available on arXiv with ID 2608.04502.
- Avoids exhaustive provisioning by using analytical methods.
Entities
Institutions
- arXiv