← alle Projekte ← all projects

cat ai-gateway.md

AI Gateway

[RUNNING]
PythonFastAPIPostgreSQLvLLMOllamaGPU OrchestrationWake-on-LAN

Der zentrale KI-Orchestrator für mein Heimnetz: eine OpenAI-kompatible Schnittstelle vor mehreren GPU-Rechnern mit lokalen Modellen — Chat, Embeddings, Transkription, Sprachausgabe, Bild- und Videogenerierung. Jedes meiner Projekte spricht ausschließlich hiermit statt mit einem Modell-Anbieter: eine Stelle für Anmeldung, eine für Begrenzung, eine für Beobachtung.

Das Gateway entscheidet, welches Modell eine Anfrage bekommt, auf welcher Karte es läuft und wer zuerst drankommt. Es weckt schlafende Rechner, wenn eine Anfrage mehr Grafikspeicher braucht als gerade wach ist, und lässt sie so lange warten. Und es zeigt seit Kurzem, wenn es überlastet ist — das konnte es lange nicht, und das war der lehrreichste Fehler des Projekts.

The central AI orchestrator for my home network: an OpenAI-compatible interface in front of several GPU machines running local models — chat, embeddings, transcription, speech, image and video generation. Every one of my projects talks to this and nothing else: one place for auth, one for limits, one for observability.

The gateway decides which model serves a request, which card it runs on and who goes first. It wakes sleeping machines when a request needs more GPU memory than is currently awake, and holds the request until then. And it recently learned to show when it is overloaded — for a long time it couldn't, and that was the most instructive bug in the project.

Familien-Chat Vorfahrt 1 Agenten-Firma Vorfahrt 7 weitere Projekte Vorfahrt 5 Gateway Modellwahl nach Faehigkeit, Last und Erfahrungswert Warteschlange mit Vorfahrt, Aufstieg gegen Verhungern Ueberlast sichtbar: Wartende, Abweisungen, Warteanteil weckt schlafende Rechner gewidmeter Motor ein Modell, dauerhaft geladen beide Karten, viele Plaetze geteilte Modelle viele, nach Bedarf geladen schlafen, wenn niemand fragt Ueberwachung Alarm aufs Handy, sobald jemand abgewiesen wurde
Ein Eingang, zwei Arten von Motor, und eine Warteschlange, die eine Reihenfolge kennt
One entrance, two kinds of engine, and a queue that knows an order

Ein Eingang statt zehn

Vorher sprach jedes Projekt direkt mit einem Modell-Server. Jedes hatte seine eigene Adresse im Code, seine eigene Fehlerbehandlung, seine eigene Vorstellung davon, was ein Zeitlimit ist. Schlief ein Rechner, fiel das dem Projekt vor die Füße.

Heute gibt es einen Eingang. Wer eine Anfrage stellt, nennt eine Fähigkeit — „Text und Code" oder „Denken" — und das Gateway entscheidet, welches Modell sie bekommt, auf welcher Karte es läuft und ob dafür erst ein Rechner geweckt werden muss. Die Anfrage wartet währenddessen einfach.

Zwei Arten von Motor, aus einem gemessenen Grund

Lange lief alles über eine Modell-Verwaltung, die Modelle nach Bedarf lädt und wieder entlädt. Das ist bequem und für gelegentliche Anfragen richtig.

Dann kam die Agenten-Firma mit vierzehn Mitarbeitern, die gleichzeitig arbeiten. Ein Vergleich beider Motorarten auf derselben Karte, mit demselben Modell, zeigte:

  • Bei ein bis zwei gleichzeitigen Anfragen gewinnt die einfache Verwaltung — sie ist in zehn Sekunden startklar.
  • Ab vier Anfragen dreht sich das Bild, und bei acht ist der andere Motor 2,6-fach schneller bei einem Bruchteil der Wartezeit.

Der Grund ist nicht Geschwindigkeit, sondern Buchhaltung: Die eine Art reserviert je Anfrage das volle Kontextfenster, die andere teilt einen gemeinsamen Speicher dynamisch.

Also beide. Ein dauerhaft geladener Motor für das eine Modell, das viele Agenten teilen — und die flexible Verwaltung für alles andere.

Die Warteschlange, die keine war

Der ältere Pfad hatte immer eine Warteschlange mit Vorfahrt: Familien-Anfragen vor Stapelarbeit, und wer lange wartet, steigt auf, damit niemand verhungert.

Der neuere Pfad hatte das nicht. Dort fragten alle gleichzeitig „ist etwas frei?", und wer im richtigen Moment fragte, gewann. Ohne Reihenfolge gibt es keine Vorfahrt — die Stufen waren wirkungslos.

Was das ausmacht, ist gemessen. Bei 136 gleichzeitig Wartenden:

  • Vorfahrt 1 kam nach 117 bis 122 Sekunden durch.
  • Vorfahrt 7 brauchte 424 Sekunden, und drei von vier Anfragen wurden abgewiesen.
Der reine Reihenfolge-Effekt ist dagegen klein, und das ist eine Selbstkorrektur: Ich hatte mit „Wartezeiten zwischen 26 und 576 Sekunden" argumentiert — aber bei 26 Anfragen pro Minute *muss* jemand zuletzt drankommen. Das ist Arithmetik, keine Ungerechtigkeit. Der Vergleich mit und ohne Reihe zeigte 0 gegen 11 Überholvorgänge von 4.560. Der Nutzen liegt in der Vorfahrt, nicht in der Reihenfolge an sich.

Der lehrreichste Fehler: Überlast war unsichtbar

Ein Lasttest mit 256 gleichzeitigen Anfragen wies 16 davon ab, mit Wartezeiten bis 118 Sekunden. Und nirgends stand etwas davon:

  • Die Anzeige „wie viele warten gerade" zählte nur die eine Sorte Warteschlange. Für den neuen Motor war sie strukturell null — 253 Stichproben mitten in der Überlast, durchgehend null.
  • Abgewiesene Anfragen bekommen keine Zeile in der Nutzungsstatistik, weil die erst nach erfolgreicher Zuteilung geschrieben wird. Während des Tests standen dort 508 Zeilen, alle mit Status „ok".
  • Das einzige eingebaute Überlast-Ereignis stand auf der Unterdrückungsliste und ging seit jeher ins Leere.

Alle drei zusammen ergaben ein System, das unter Volllast „gesund" meldete. Die Gesundheitsprüfung tat dasselbe: Sie fasst nur die Datenbank an und antwortet mit 200, egal ob eine Antwort zwei Sekunden oder zehn Minuten braucht.

Seither gibt es eine ehrliche Kennzahl — wie viele warten, wie viele wurden abgewiesen, welcher Anteil musste überhaupt warten — und einen Alarm aufs Handy, sobald jemand abgewiesen wird. Es ist die erste Überwachung im Haus, die nicht fragt „läuft es noch?", sondern „reicht es noch?".

Wo die Grenze liegt

Gemessen, nicht geschätzt: rund 26 Anfragen pro Minute. Von 16 auf 256 gleichzeitige Anfragen verachtfacht sich die Last, der Durchsatz steigt um 43 Prozent und bleibt dann stehen. Alles darüber wird nur noch in Wartezeit umgewandelt.

Die Plätze am Motor zu verdoppeln brachte nichts. Die Grenze ist die Hardware, nicht die Einstellung — und das zu wissen ist mehr wert als jede Vermutung darüber.

One entrance instead of ten

Before, every project talked to a model server directly. Each had its own address in the code, its own error handling, its own idea of what a timeout is. When a machine was asleep, the project had to deal with it.

Today there is one entrance. A caller names a capability — "text and code" or "reasoning" — and the gateway decides which model serves it, on which card it runs, and whether a machine has to be woken first. The request simply waits.

Two kinds of engine, for a measured reason

For a long time everything ran through a model manager that loads and unloads models on demand. That is convenient and right for occasional requests.

Then came the agent company with fourteen employees working at the same time. Comparing both kinds of engine on the same card, with the same model, showed:

  • At one or two concurrent requests the simple manager wins — it is ready in ten seconds.
  • From four requests the picture flips, and at eight the other engine is 2.6× faster at a fraction of the waiting time.

The reason isn't speed, it's bookkeeping: one kind reserves the full context window per request, the other shares a pool dynamically.

So: both. A permanently loaded engine for the one model many agents share — and the flexible manager for everything else.

The queue that wasn't one

The older path always had a queue with priorities: family requests before batch work, and whoever waits long moves up so nobody starves.

The newer path had none. There, everyone asked "is anything free?" at the same time, and whoever asked at the right moment won. Without an order there is no priority — the levels had no effect.

What that costs is measured. With 136 waiting at once:

  • Priority 1 got through after 117 to 122 seconds.
  • Priority 7 took 424 seconds, and three of four requests were rejected.
The pure ordering effect, by contrast, is small — and that is a self-correction: I had argued with "waiting times between 26 and 576 seconds", but at 26 requests per minute somebody *must* come last. That is arithmetic, not unfairness. The comparison with and without the queue showed 0 against 11 overtakes out of 4,560. The value is in the priority, not in the ordering itself.

The most instructive bug: overload was invisible

A load test with 256 concurrent requests rejected 16 of them, with waits up to 118 seconds. And nowhere was any of it visible:

  • The "how many are waiting" figure only counted one kind of queue. For the new engine it was structurally zero — 253 samples taken during the overload, zero throughout.
  • Rejected requests get no row in the usage log, because that is only written after a successful allocation. During the test it held 508 rows, all with status "ok".
  • The one built-in overload event was on the suppression list and had always gone nowhere.

Together these produced a system that reported "healthy" at full load. The health check did the same: it only touches the database and answers 200 whether a reply takes two seconds or ten minutes.

There is now an honest metric — how many are waiting, how many were rejected, what share had to wait at all — plus an alert to my phone as soon as someone is rejected. It is the first monitor in the house that doesn't ask "is it still running?" but "is it still enough?".

Where the limit is

Measured, not estimated: about 26 requests per minute. Going from 16 to 256 concurrent requests multiplies the load eightfold, throughput rises by 43 percent and then stops. Everything beyond that is only converted into waiting time.

Doubling the slots at the engine changed nothing. The limit is the hardware, not the configuration — and knowing that is worth more than any guess about it.

18