Latest

Doctoral Researcher — ICCS, National Technical University of Athens

Om Rajendra
Kathalkar

I make large language models remember less and answer just as well — efficient inference for real, resource-constrained deployments.

Om Kathalkar on a rooftop in Athens at sunset
Zografou, Athens · 2026
Role
Doctoral Researcher, ICCS / NTUA
Project
f-inference
Advisor
Prof. Konstantinos Tserpes ↗
Based in
Athens, Greece
My research in 20 seconds

What is KV-cache compression?

While an LLM reads, it stores a key and a value for every token. On long prompts and long reasoning chains that cache fills GPU memory. Eviction keeps the tokens that matter and drops the rest. Drag the budget to see it.

Tokens kept

Illustration only — importance scores are hand-set for this demo. The first token is always kept: models lean on it as an attention sink.

[BOS]Toansweralongquestion,themodelstoresakeyandvalueforeverytokenithasread.Mostofthembarelymatter,sowekeepthefewthatdoandfreetherest.
01

About

I'm a PhD researcher at the Institute of Communication and Computer Systems (ICCS), National Technical University of Athens, working on efficient AI inference for resource-constrained and distributed environments.

CurrentlyDoctoral Researcher, ICCS, National Technical University of Athens
Projectf-inference — Efficient Foundation Model Inference Across the Computing Continuum
FocusEfficient inference for large language and foundation models
PreviouslyResearcher, Ericsson Research AI · MS by Research, IIIT Hyderabad
LanguagesEnglish, Marathi, Hindi

My research sits where systems meet machine learning: how large language models can be made leaner in what they store, move and compute, without losing the reasoning quality that makes them useful.

Before Athens I completed an MS by Research at IIIT Hyderabad, building TRAQID, a 26,678-image traffic air-quality dataset, and AQIFormer, a transformer that classifies AQI across cities. At Ericsson Research AI I worked on network-aware path planning for factory robots in the Horizon Europe PANDORA project.

I believe inference efficiency is the bottleneck between lab-scale models and real deployment, and the most interesting work lives at that boundary. Outside research: cricket, hockey, badminton, hiking, and writing short stories and poetry.

02

Research

Current project

f-inference

Efficient Foundation Model Inference Across the Computing Continuum

Foundation models are trained on broad data and adapt to tasks across text, images, audio, video and time series — but running them today depends on centralised cloud providers, costs a lot of energy and money, and adds latency.

f-inference builds a Resource-driven Computing Continuum that distributes foundation-model inference across Cloud, Edge and IoT, moving it closer to data and users through adaptive, efficient inference — cutting latency, energy and cost while preserving data sovereignty, and opening advanced GenAI to SMEs and resource-constrained communities.

Challenges it tackles

  1. 1Dependence on centralised cloud providers, threatening digital sovereignty and data security
  2. 2High energy demands that conflict with sustainability goals
  3. 3Costs that exclude SMEs from deploying and customising models
  4. 4Remote services that add latency and single points of failure
01Efficient LLM inferenceEfficient inference for large language models2025 — now

Large language models are powerful but costly to run: their sheer size, attention that grows quadratically with input length, and token-by-token decoding drive up latency, memory and energy. My research makes LLM inference leaner across the data, model and system levels — reducing what a model stores, moves and computes — so it runs fast and reliably on constrained hardware, closer to where the data lives.

02Robotics · NetworksNetwork-aware path planning for industrial robots2024 — 2025

Factory robots are limited by wireless quality as much as by floor space. With Ericsson Research AI I built a Sionna ray-tracing framework validated on real measurements (R² = 0.87) and planners that co-optimise path cost and coverage; a spatial-aware transformer predicts SNR within 2.31 dB and runs 31× faster than GNN baselines.

  • AMR path planning
  • Sionna ray tracing
  • CVAE · GraphMP
  • IEEE ANTS 2025
  • Springer Autonomous Robots
03Computer vision · Air qualityVisual air-quality intelligence2021 — 2025

Estimating air quality from street-level images removes the dependency on expensive sensor networks. I built TRAQID — 26,678 front and rear traffic images with co-located AQI and weather data — and AQIFormer, a multi-view transformer reaching 89.96% accuracy and 81.67% on a city unseen in training.

  • TRAQID
  • AQIFormer
  • Vision transformers
  • Cross-city generalisation
  • PAN-AQI
04IoT · EducationRemote, vision-evaluated laboratories2021 — 2022

A platform that lets students run real physics experiments over the internet, with computer vision measuring the outcome — about 10× lower error than infrared sensing. It served thousands of learners and led to a FiCloud 2022 paper and a US patent.

  • Remote labs
  • Computer vision
  • IoT
  • FiCloud 2022
  • US patent
  1. Nov 2025 — nowDoctoral ResearcherICCS, NTUA · AthensEfficient foundation-model inference, f-inference project.
  2. Dec 2024 — Nov 2025ResearcherEricsson Research AINetwork-aware AMR path planning, Horizon Europe PANDORA.
  3. Apr 2024 — nowAI/ML MentorTalentSprint (Accenture)AI Infinity and AI/ML certification programmes.
  • Sep 2026PAN-AQI accepted at the IEEE World Forum on Internet of Things (WFIoT) 2026.
  • Nov 2025Joined ICCS / NTUA as a Doctoral Researcher.
  • Nov 2025Defended my MS by Research thesis, "Camera-Based Deep Learning Framework for AQI Estimation: Dataset and Methodology", at IIIT Hyderabad.
07

Contact

The best way to reach me is email. Happy to talk research, collaborations, or a good idea.

Personal: omkathalkar.connect@gmail.com

School of Electrical and Computer Engineering
National Technical University of Athens
Polytechnic Campus, Zografou · 157 73 Athens, Greece