其他行业 · AI应用
使用pplx-api实现高效LLM推理
Perplexity提供pplx-api,集成开源大语言模型,利用NVIDIA TensorRT-LLM实现快速推理。该方案运行在配备NVIDIA A100 GPU的AWS EC2 P4d实例上,并正在迁移至配备NVIDIA H100 GPU的P5实例。旨在优化成本同时保持实时应用的高性能。
推理延迟降低(相对于其他部署平台)
up to 3.1X lower latency
来源披露
首个Token延迟降低(相对于其他部署平台)
up to 4.3X lower first-token latency
来源披露
成本节省
$600,000 per year
来源披露
01
业务背景
Perplexity提供pplx-api,集成开源大语言模型,利用NVIDIA TensorRT-LLM实现快速推理。该方案运行在配备NVIDIA A100 GPU的AWS EC2 P4d实例上,并正在迁移至配备NVIDIA H100 GPU的P5实例。旨在优化成本同时保持实时应用的高性能。
02
遇到的问题
Perplexity faces challenges managing escalating costs of LLM inference to support rapid growth, optimizing infrastructure for massive scale at minimal cost while meeting strict SLA requirements, and adapting quickly to explosive growth of community LLMs that drives up cost and deployment complexity.
03
AI 解决方案
Perplexity deploys pplx-api on Amazon EC2 P4d instances with NVIDIA A100 GPUs, accelerated by NVIDIA TensorRT-LLM for optimized inference. They plan full transition to Amazon P5 instances with NVIDIA H100 GPUs to further cut latency and boost throughput. Software optimizations use TensorRT-LLM for FlashAttention and masked MHA. For scalability, they use AWS integration with Kubernetes to elastically scale beyond hundreds of GPUs.
04
实施结果
推理延迟降低(相对于其他部署平台):up to 3.1X lower latency(up to 3.1X lower)
首个Token延迟降低(相对于其他部署平台):up to 4.3X lower first-token latency(up to 4.3X lower)
成本节省:$600,000 per year(4X cost reduction)
迁移至H100后的延迟降低(相对于A100):cut latency in half(50% reduction)
迁移至H100后的吞吐量提升(相对于A100):boost throughput by 200 percent(200% increase)
05
风险与边界
未披露
信息来源
1 个来源页面内容为结构化改写;关键结论应可追溯到以下材料。