NVIDIA

NVIDIA Triton Inference Server

NVIDIA Dynamo-Triton (formerly NVIDIA Triton Inference Server) is open-source software that enables deployment of AI models across major frameworks including TensorRT, PyTorch, ONNX, OpenVINO, Python and RAPIDS FIL, with dynamic batching and concurrent execution, supporting real-time, batched, ensemble and audio/video streaming workloads on NVIDIA GPUs, non-NVIDIA accelerators, x86 and ARM CPUs.

Official product page
2published use cases
2industries
1countries on record
Large Language Models
top AI capability

Evidence mix: High 2 · Medium 0 · Low 0 — bands are computed from each record's evidence signals.

Industry
Country

2 use cases

Aerospace & DefenseGenerative AILarge Language ModelsRetrieval-Augmented GenerationAI Model Development & MLOps

Boosting Innovation and Cutting Costs Through Lockheed Martin's AI Factory

Lockheed Martin centralized compute resources, MLOps tools and best practices into the Lockheed Martin AI Factory, built on an NVIDIA DGX SuperPOD reference architecture, to build and deploy trustworthy AI at scale on-premises under strict data governance requirements. Developers can now get GPU-backed environments running in minutes instead of weeks, and training times dropped from weeks to days. The AI factory processes over one billion tokens per week and now serves 7,000 engineers and developers, supporting internal chatbots and coding assistants such as Lockheed Martin Text Navigator, and has consolidated 30+ models on-premises.

Lockheed MartinNVIDIA DGX SuperPOD · NVIDIA Triton Inference Server · NVIDIA NeMo Framework +2