f-inference
Efficient Foundation Model Inference Across the Computing Continuum
Foundation models are trained on broad data and adapt to tasks across text, images, audio, video and time series — but running them today depends on centralised cloud providers, costs a lot of energy and money, and adds latency.
f-inference builds a Resource-driven Computing Continuum that distributes foundation-model inference across Cloud, Edge and IoT, moving it closer to data and users through adaptive, efficient inference — cutting latency, energy and cost while preserving data sovereignty, and opening advanced GenAI to SMEs and resource-constrained communities.
Challenges it tackles
- 1Dependence on centralised cloud providers, threatening digital sovereignty and data security
- 2High energy demands that conflict with sustainability goals
- 3Costs that exclude SMEs from deploying and customising models
- 4Remote services that add latency and single points of failure