Technology & SoftwareGenerative AIPublic Cloud

Accelerating Large Language Model Inference with NVIDIA in the Cloud

PerplexityNVIDIA TensorRT-LLM · NVIDIA A100 Tensor Core GPUs · NVIDIA H100 Tensor Core GPUs +2

Perplexity built pplx-api, an API for developers to integrate open-source LLMs with fast inference, served on Amazon EC2 P4d instances powered by NVIDIA A100 Tensor Core GPUs and accelerated with NVIDIA TensorRT-LLM (with a planned move to Amazon P5 instances with NVIDIA H100 GPUs). pplx-api achieves up to 3.1X lower latency and up to 4.3X lower first-token latency versus other deployment platforms, and switching external inference-serving API references to pplx-api lowered costs 4X, saving $600,000 per year. Using NVIDIA H100 GPUs and FP8 precision on Amazon P5 instances cuts latency in half and boosts throughput by 200 percent versus A100 GPUs in the same configuration. Perplexity also uses AWS's Kubernetes integration to scale elastically beyond hundreds of GPUs.

Overview

Perplexity built pplx-api, an API for developers to integrate open-source LLMs with fast inference, served on Amazon EC2 P4d instances powered by NVIDIA A100 Tensor Core GPUs and accelerated with NVIDIA TensorRT-LLM (with a planned move to Amazon P5 instances with NVIDIA H100 GPUs). pplx-api achieves up to 3.1X lower latency and up to 4.3X lower first-token latency versus other deployment platforms, and switching external inference-serving API references to pplx-api lowered costs 4X, saving $600,000 per year. Using NVIDIA H100 GPUs and FP8 precision on Amazon P5 instances cuts latency in half and boosts throughput by 200 percent versus A100 GPUs in the same configuration. Perplexity also uses AWS's Kubernetes integration to scale elastically beyond hundreds of GPUs.

The challenge

Delivering fast and efficient LLM inference is critical for real-time applications. As a startup, Perplexity faced escalating costs associated with LLM inference to support its rapid growth, while needing to maintain strict service-level agreement requirements and adapt quickly to an explosively growing ecosystem of community LLMs.

The solution

Perplexity built pplx-api, an API for developers to integrate open-source LLMs with fast inference, served on Amazon EC2 P4d instances powered by NVIDIA A100 Tensor Core GPUs and accelerated with NVIDIA TensorRT-LLM, with a planned full transition to Amazon P5 instances powered by NVIDIA H100 Tensor Core GPUs. Perplexity uses AWS's integration with Kubernetes to scale elastically beyond hundreds of GPUs.

Generative AILarge Language Models

Reported business value

pplx-api achieves up to 3.1X lower latency and up to 4.3X lower first-token latency relative to other deployment platforms. Switching external inference-serving API references to pplx-api lowered costs 4X, saving $600,000 per year. Using NVIDIA H100 GPUs and FP8 precision on Amazon P5 instances cuts latency in half and boosts throughput by 200 percent compared to NVIDIA A100 GPUs in the same configuration.

Sources

Open any source and check the claim yourself — that is the point of the register.

This record was researched and written with AI assistance, and its claims were checked against the sources above. (EU AI Act art. 50 transparency notice.)

Related entries

Other technology & software entries in the register.

All entries
Technology & SoftwareRecommendation & PersonalizationPublic Cloud

HP crafts marketing campaigns that resonate with customers using Databricks and Uniphore

HP centralized first-party customer data on the Databricks Data + AI Platform with Delta Lake and Unity Catalog, and connected it to Uniphore's HybridCompute for federated query pushdown, cutting campaign setup from 2 weeks to 2 hours and processing 400 million records in seconds.

96/100HighPrimary source
HPDelta Lake · Unity Catalog
Technology & SoftwarePredictive AnalyticsUnknown

Transforming Weather Forecasting with Lakeflow Jobs

AccuWeather migrated from on-premises infrastructure to Databricks and Lakeflow Jobs, working with Datadog for observability, to unify diverse weather data formats and orchestrate 4,500+ weekly jobs. Lakeflow Jobs coordinates the ingestion of multiple weather models, triggers machine learning processes that weight and blend different forecasts, and manages complex job dependencies for reinforcement training workflows used in AccuWeather's proprietary forecasting engine. AccuWeather reports 3x faster dataset development (three months to one month per dataset), a 50% reduction in unactionable alerts, and 50% cost savings on serverless job usage.

92/100HighPrimary source
AccuWeatherLakeflow Jobs · Delta Sharing · Datadog
Technology & SoftwareRecommendation & PersonalizationPublic Cloud

Adobe brings creativity to life with Databricks

Adobe uses the Databricks Data + AI Platform for end-to-end data management that unifies all data and AI at scale, with 20% faster performance. Databricks equips over 92 teams at Adobe to unify data from financials, sales, products, customers and employees so they can drive personalized experiences across Adobe's digital platforms with AI.

92/100HighPrimary source
Adobe
Technology & SoftwareAgentic AIPublic Cloud

Supermetrics: Helping Marketers Redefine Efficiency with AI-Powered Data Analysis

Supermetrics, a Finland-based marketing intelligence platform serving 15,000+ customers across 132 countries, built an AI agent on Google Cloud using Vertex AI Agent Builder and the Agent Development Kit (ADK) that autonomously manages data connections, fixes pipeline errors, and analyzes campaign performance in real time, suggesting new creative options using Imagen. The agent automates the weekly marketing reporting cycle that previously took performance marketers up to four hours, reclaiming over 15 hours per month per marketer for strategy and creative testing. The system uses a central AI agent that interprets natural language requests and delegates tasks to sub-agents, and stores 'core memories' of user preferences for personalized context.

96/100HighPrimary source
Supermetrics· FinlandGoogle Cloud · Vertex AI Agent Builder · Agent Development Kit (ADK) +1

Was this helpful?

Your feedback helps us improve our use case database