Accelerating Large Language Model Inference with NVIDIA in the Cloud
Perplexity built pplx-api, an API for developers to integrate open-source LLMs with fast inference, served on Amazon EC2 P4d instances powered by NVIDIA A100 Tensor Core GPUs and accelerated with NVIDIA TensorRT-LLM (with a planned move to Amazon P5 instances with NVIDIA H100 GPUs). pplx-api achieves up to 3.1X lower latency and up to 4.3X lower first-token latency versus other deployment platforms, and switching external inference-serving API references to pplx-api lowered costs 4X, saving $600,000 per year. Using NVIDIA H100 GPUs and FP8 precision on Amazon P5 instances cuts latency in half and boosts throughput by 200 percent versus A100 GPUs in the same configuration. Perplexity also uses AWS's Kubernetes integration to scale elastically beyond hundreds of GPUs.
Overview
Perplexity built pplx-api, an API for developers to integrate open-source LLMs with fast inference, served on Amazon EC2 P4d instances powered by NVIDIA A100 Tensor Core GPUs and accelerated with NVIDIA TensorRT-LLM (with a planned move to Amazon P5 instances with NVIDIA H100 GPUs). pplx-api achieves up to 3.1X lower latency and up to 4.3X lower first-token latency versus other deployment platforms, and switching external inference-serving API references to pplx-api lowered costs 4X, saving $600,000 per year. Using NVIDIA H100 GPUs and FP8 precision on Amazon P5 instances cuts latency in half and boosts throughput by 200 percent versus A100 GPUs in the same configuration. Perplexity also uses AWS's Kubernetes integration to scale elastically beyond hundreds of GPUs.
The challenge
Delivering fast and efficient LLM inference is critical for real-time applications. As a startup, Perplexity faced escalating costs associated with LLM inference to support its rapid growth, while needing to maintain strict service-level agreement requirements and adapt quickly to an explosively growing ecosystem of community LLMs.
The solution
Perplexity built pplx-api, an API for developers to integrate open-source LLMs with fast inference, served on Amazon EC2 P4d instances powered by NVIDIA A100 Tensor Core GPUs and accelerated with NVIDIA TensorRT-LLM, with a planned full transition to Amazon P5 instances powered by NVIDIA H100 Tensor Core GPUs. Perplexity uses AWS's integration with Kubernetes to scale elastically beyond hundreds of GPUs.
Reported business value
pplx-api achieves up to 3.1X lower latency and up to 4.3X lower first-token latency relative to other deployment platforms. Switching external inference-serving API references to pplx-api lowered costs 4X, saving $600,000 per year. Using NVIDIA H100 GPUs and FP8 precision on Amazon P5 instances cuts latency in half and boosts throughput by 200 percent compared to NVIDIA A100 GPUs in the same configuration.
Sources
Open any source and check the claim yourself — that is the point of the register.
This record was researched and written with AI assistance, and its claims were checked against the sources above. (EU AI Act art. 50 transparency notice.)
Other technology & software entries in the register.
HP crafts marketing campaigns that resonate with customers using Databricks and Uniphore
HP centralized first-party customer data on the Databricks Data + AI Platform with Delta Lake and Unity Catalog, and connected it to Uniphore's HybridCompute for federated query pushdown, cutting campaign setup from 2 weeks to 2 hours and processing 400 million records in seconds.
Transforming Weather Forecasting with Lakeflow Jobs
AccuWeather migrated from on-premises infrastructure to Databricks and Lakeflow Jobs, working with Datadog for observability, to unify diverse weather data formats and orchestrate 4,500+ weekly jobs. Lakeflow Jobs coordinates the ingestion of multiple weather models, triggers machine learning processes that weight and blend different forecasts, and manages complex job dependencies for reinforcement training workflows used in AccuWeather's proprietary forecasting engine. AccuWeather reports 3x faster dataset development (three months to one month per dataset), a 50% reduction in unactionable alerts, and 50% cost savings on serverless job usage.
Adobe brings creativity to life with Databricks
Adobe uses the Databricks Data + AI Platform for end-to-end data management that unifies all data and AI at scale, with 20% faster performance. Databricks equips over 92 teams at Adobe to unify data from financials, sales, products, customers and employees so they can drive personalized experiences across Adobe's digital platforms with AI.
Supermetrics: Helping Marketers Redefine Efficiency with AI-Powered Data Analysis
Supermetrics, a Finland-based marketing intelligence platform serving 15,000+ customers across 132 countries, built an AI agent on Google Cloud using Vertex AI Agent Builder and the Agent Development Kit (ADK) that autonomously manages data connections, fixes pipeline errors, and analyzes campaign performance in real time, suggesting new creative options using Imagen. The agent automates the weekly marketing reporting cycle that previously took performance marketers up to four hours, reclaiming over 15 hours per month per marketer for strategy and creative testing. The system uses a central AI agent that interprets natural language requests and delegates tasks to sub-agents, and stores 'core memories' of user preferences for personalized context.
Was this helpful?
Your feedback helps us improve our use case database

