IN
Kedar Sathe Avatar

KEDAR SATHE

@wtfkedar

I'm Kedar, a Machine Learning Engineer from Pune. I build end-to-end ML systems, deep learning models, and scalable MLOps infrastructure.

Most of my time goes into model training, optimization, and inference pipelines. As an OSS Maintainer of mermaid-js, and a contributor to projects like vLLM and Weaviate, I enjoy improving serving efficiency and developer tooling.

I also design agentic workflows and retrieval-augmented systems (RAG), and on the systems side, I enjoy low-level programming like custom kernel modules and performance tuning.

Open to ML engineering roles, collaborations, and technical discussions — feel free to reach out.

Where you can Connect with Me (of course digitally)

Work Experience

Autonex AI 360 Logo

Machine Learning Engineer

Autonex AI 360

  • Fine-tuned and optimized domain-specific LLMs using PEFT techniques (LoRA, QLoRA)
  • Integrated OpenCV computer vision algorithms into live camera streams for real-time video processing
  • Worked on the internal PM (Project Management) portal to track model training runs and visualization metrics
  • Dockerized inference workloads and benchmarked latency across various hardware targets
PyTorchOpenCVPEFTReactDockerPython

Jun 2026 – Present

Remote · Mumbai, India

Mermaid.js Logo

Open-Source Maintainer

Mermaid.js

  • Maintains Mermaid.js, a widely used open-source diagramming and visualization library
  • Triages issues, reviews pull requests, and guides open-source contributors globally
  • Resolved 8+ bugs affecting rendering fidelity and cross-platform compatibility
  • Shipped PRs with new functionality and parser optimizations
JavaScriptTypeScriptD3.jsSVGGitHubOpen Source

Jun 2025 – Present

Remote

Unique School App Logo

Software Developer Intern (ML & Backend)

Unique School App India

  • Developed a multi-agent RAG system using LangGraph to enable context-aware Q&A over school administrative data
  • Designed routing, hybrid retrieval, and response synthesis pipelines for administrators
  • Led React Native app migration to integrate AI-powered features and cross-platform UI components
  • Set up and improved CI/CD DevOps pipelines for automated backend and API deployment
React NativeLangGraphRAGREST APIDevOpsNode.js

Jan 2026 – Apr 2026

Pune, India

ML & MLOps Engineer

Freelance

  • Built responsive web applications integrated with custom LLMs, vector search, and hybrid RAG systems
  • Implemented end-to-end ML training and evaluation pipelines, deploying models on cloud platforms
  • Worked with clients to deliver stable, scalable systems containerized with Docker
Next.jsReactPythonTensorFlowDockerAzure

2023 – Present

Remote

Tech Stack

This list grows faster than my training loss curves — and I love that.

< Machine Learning & Deep Learning />

PyTorch
TensorFlow
Scikit-Learn
Hugging Face
PEFT / LoRA
Unsloth
Deep Q-Learning

< MLOps & Infrastructure />

Docker
Kubernetes
vLLM
Triton Inference Server
AWS
Google Cloud
Linux
Git
GitHub

< Languages />

Python
TypeScript
JavaScript
C / C++

< Data & Databases />

PostgreSQL
pgvector
MongoDB
Apache Spark

< Frameworks & Web Dev />

LangChain
FastAPI
React
Next.js
Node.js
TailwindCSS

< Developer Tools />

VS Code
Vercel
Postman
Figma

Proof of Work

Things I built when curiosity got the better of me.

Small-LLM-RLHF
Live

Small-LLM-RLHF

From-scratch, educational implementation of 3D-parallel RLHF training for small language models. Built on PyTorch and Triton with custom FlashAttention kernels, PagedAttention engine, and async PPO loop.

PyTorch
Triton
RLHF
3D Parallelism
PagedAttention
GroundedRAG
Live

GroundedRAG

Local multi-agent RAG system built with LangGraph, Ollama, and pgvector. Designed to stay grounded, refuse weak answers, and show its retrieval trail using hybrid search + RRF fusion.

LangGraph
Ollama
pgvector
Streamlit
Python
LazyDocs
Live

LazyDocs

CLI-based NPM module that automates documentation generation using the GroQ API. Analyzes project structure and generates structured docs — designed for terminal-based developer workflows.

Node.js
CLI
NPM
GroQ API
AI
PitchForge
Live

PitchForge

AI-powered outreach generator built with Next.js & Tailwind CSS v4. Generates tailored cold emails, cover letters, and resume suggestions using Gemini/Groq, with coordinate-based PDF parsing.

Next.js
TailwindCss
Gemini
Groq
AI
Pelvix AI
Live

Pelvix AI

Multimodal AI chatbot with personality. Supports text and image generation, real-time conversational responses via DeepSeek-R1 and Flux. Privacy-first: session-only chats, no tracking.

Next.js
React
DeepSeek-R1
GSAP
AI
X-Bird
Live

X-Bird

Chrome extension for AI-generated contextual replies on X.com. Supports multiple tones: Auto, Savage, Supportive, Funny, Professional, Casual — with adjustable intensity control.

Chrome Extension
AI
Hugging Face
Qwen2.5-72B
Snake's & The Golden Apple
Live

Snake's & The Golden Apple

AI-driven competitive snake game using Deep Q-Learning. Multi-agent gameplay with Double DQN, Dueling DQN, and Prioritized Experience Replay. 25D state space, 50K replay buffer.

Python
PyTorch
Deep Q-Learning
RL
Pygame
FreeLix
Live

FreeLix

Free transcription and translation application powered by Whisper-large-v3. Provides audio transcription and multilingual translation — fully open-source and accessible.

React
Whisper
Hugging Face
Tailwind CSS
AI

Blogs & Notes

I write about AI experiments, system design breakdowns, and lessons from building in public. Check out my active blog for technical deep-dives and dev notes.

neuronreads.vercel.app

AI, ML systems, dev notes & more →