Harmonia: End-to-End RAG Serving Optimization

arXiv CS Tuesday 09 June 2026, 04:00 UTC By Saurabh Agarwal, Bodun Hu, Luis Pabon, Myungjin Lee, Jayanth Srinivasa, Aditya Akella 1 min read

Key Points

arXiv:2505.07833v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) improves the reliability of large language models by integrating external knowledge, but serving RAG pipelines efficiently is challenging because requests traverse heterogeneous components spanning LLM inference, databases, and CPU-side processing. We present Harmonia, an end-to-end RAG serving framework that addresses these bottlenecks through (i) a flexible pipeline specification interface for composing custom workflows, (ii) heterogeneity-aware deployment that provisions and configures components as a distributed inference system, and (iii) a closed-loop runtime controller that monitors load and execution progress and reduces SLO violations through request prioritization and auto-scaling. Across four RAG applications, Harmonia outperforms commercial alternatives, improving throughput by more than 2.04x while reducing SLO violations by up to 78.4 percent.

LLM (ORG) CPU (ORG) SLO (ORG)

Originally published by arXiv CS Read original →

Nasa chief defends choice of all-male Artemis III crew Critics fear the agency is following Trump’s order to eliminate diversity and inclusion efforts despite its vow to put a woman on the moon Nasa’s administrator Jared Isaacman on Wednesday defended the make-up of the space agency’s latest Artemis crew, an all-male group. The nominations have earned criticism that Nasa may have acted in accordance with US President Donald Trump’s direction to eliminate diversity and inclusion efforts....

South China Morning Post 20m ago

The asteroid that wiped out the dinosaurs may have created a vast underground habitat for life that lasted 8 million years

The asteroid that wiped out the dinosaurs may have created a vast underground habitat for life that lasted 8 million years The Chicxulub impact may have actually helped nurture life while destroying it, too. The asteroid impact that doomed the dinosaurs may also have built one of Earth's longest-lasting underground ecosystems. When a roughly 6-mile-wide (10-kilometer-wide) asteroid slammed into what is now Mexico's Yucatán Peninsula 66 million years ago, it triggered a global catastrophe...

Space.com 22m ago

See the 'crawling,' ball-shaped robot that rolled around the moon during Japan's historic first landing

See the 'crawling,' ball-shaped robot that rolled around the moon during Japan's historic first landing A morphable moon robot operated for 100 minutes in 2024, allowing investigators to get images of an upside-down spacecraft on the lunar surface. When the Japanese Smart Lander for Investigating Moon (SLIM) spacecraft, nicknamed the "Moon Sniper," face-planted onto the lunar surface in 2024, an experimental rover told Earth scientists what happened. Rolling autonomously through the lunar...

Live Science 22m ago

Harmonia: End-to-End RAG Serving Optimization

Related Stories

'Worrying' pollution in Cotswolds river - volunteers

Nasa chief defends choice of all-male Artemis III crew

The asteroid that wiped out the dinosaurs may have created a vast underground habitat for life that lasted 8 million years

See the 'crawling,' ball-shaped robot that rolled around the moon during Japan's historic first landing