LogisticsLarge Language Models
Delhivery achieves 160 ms latency for high-precision geocoding using Amazon EKS
Delhivery, a logistics provider in India, implemented a fine-tuned open-source Llama 3.2 1B large language model on Amazon EKS to support high-volume geocoding of pickup and drop-off addresses. The system processes up to 8,000 requests per minute at 160 milliseconds latency using NVIDIA A10G GPU-backed G5 Xlarge instances and the vLLM framework. Delhivery cut model-serving costs by approximately 80 percent and accelerated prototyping cycles from two days to under six hours, working with the AWS Prototyping and Cloud Engineering (PACE) team.
Delhivery· IndiaLlama 3.2 1B · Amazon Elastic Kubernetes Service · Amazon EC2 G5 Instances +3