Architecting and serving high-performance, containerized multi-modal generative AI pipelines.
Clear code, simple architecture, and machine learning that actually works.
Fast LLM serving with vLLM, speech synthesis (TTS), and multi-modal pipelines optimized for low latency.
Async APIs built with FastAPI, containerized with Docker, running on Ubuntu Linux GPU servers.
I am an AI Engineer from India. I write Python backends, optimize model serving, and build multi-modal AI pipelines.
Currently working as an AI Engineer Intern at Anivale. My focus is on vLLM, speech synthesis (TTS), video generation, FastAPI, and Docker Compose on remote GPU servers.
I believe the best software is simple. I focus on clarity, performance, and real utility over complexity.
What I do day-to-day.
AI Engineer Intern · Internship
Architecting and serving high-performance, containerized multi-modal generative AI pipelines.
Orchestrating production-ready data streams across advanced LLMs, low-latency speech synthesis (TTS), and automated video generation engines.
Implementing vLLM and vLLM-Omni to maximize model throughput, minimize token latency, and optimize remote GPU resource allocation.
Engineering highly concurrent, asynchronous backend services using FastAPI and standardizing deployments across multi-container Docker Compose environments.
Active AI systems and model serving pipelines under development.
Currently working on production multi-modal generative AI pipelines, low-latency speech synthesis (TTS), and high-throughput vLLM model serving infrastructure at Anivale. Detailed project write-ups and open-source code repositories will be published here soon.
Foundational software coursework.
University of Helsinki • Dec 2024
Production-grade curriculum covering JavaScript, React, Node.js, Express.js, MongoDB, REST APIs, TypeScript, CI/CD pipelines, and Docker containers.
View CertificateI'm based in India and open for remote AI engineering roles, high-throughput model serving infrastructure, and multi-modal pipeline architecture.