
  <rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
      <title>KI im Mittelstand – News &amp; Technologie-Übersicht</title>
      <link>https://www.ki-mittelstand.eu/blog</link>
      <description>Faktenbasierte News, Technologien und Praxislösungen zu Künstlicher Intelligenz für mittelständische Unternehmen. Ohne Hype, mit klaren Quellen.</description>
      <language>de-de</language>
      <managingEditor>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</managingEditor>
      <webMaster>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</webMaster>
      <lastBuildDate>Tue, 04 Aug 2026 00:00:00 GMT</lastBuildDate>
      <atom:link href="https://www.ki-mittelstand.eu/tags/vllm/feed.xml" rel="self" type="application/rss+xml"/>
      
  <item>
    <guid>https://www.ki-mittelstand.eu/blog/deepseek-v4-flash-eigene-hardware</guid>
    <title>DeepSeek V4 Flash selbst hosten: VRAM, Setup, Durchsatz</title>
    <link>https://www.ki-mittelstand.eu/blog/deepseek-v4-flash-eigene-hardware</link>
    <description>284 Milliarden Parameter unter MIT-Lizenz. Was der Betrieb im eigenen Rechenzentrum wirklich kostet — Speicherrechnung, Konfiguration, Grenzen.</description>
    <pubDate>Tue, 04 Aug 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>deepseek</category><category>open-weights</category><category>vllm</category><category>gpu</category><category>on-premise</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/ki-server-api-design-openai-kompatibel</guid>
    <title>KI-Server-API-Design: OpenAI-kompatible APIs im Unternehmen</title>
    <link>https://www.ki-mittelstand.eu/blog/ki-server-api-design-openai-kompatibel</link>
    <description>OpenAI-kompatible APIs für den KI-Server: vLLM, Ollama und LiteLLM im Vergleich — API-Design, Authentifizierung und Enterprise-Integration.</description>
    <pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>api</category><category>openai-compatibel</category><category>vllm</category><category>ollama</category><category>ki-server</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/ollama-vs-vllm-vs-lm-studio-llm-frontendvergleich</guid>
    <title>Ollama vs vLLM vs LM Studio: Welches LLM-Frontend für den Mittelstand?</title>
    <link>https://www.ki-mittelstand.eu/blog/ollama-vs-vllm-vs-lm-studio-llm-frontendvergleich</link>
    <description>Ollama, vLLM und LM Studio im Vergleich: Welches LLM-Frontend passt für den deutschen Mittelstand? Performance, Features, Deployment und Kosten.</description>
    <pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ollama</category><category>vllm</category><category>lm-studio</category><category>llm-frontends</category><category>self-hosted</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/ollama-zu-vllm-migration-produktionsreife</guid>
    <title>Von Ollama zu vLLM: wann der Pilot produktiv werden muss</title>
    <link>https://www.ki-mittelstand.eu/blog/ollama-zu-vllm-migration-produktionsreife</link>
    <description>Ollama trägt den Piloten, nicht den Betrieb. Die vier Signale für den Wechsel auf vLLM — und die Migration Schritt für Schritt.</description>
    <pubDate>Tue, 04 Aug 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ollama</category><category>vllm</category><category>migration</category><category>inferenz</category><category>produktivbetrieb</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/strukturierte-llm-ausgaben-json-schema</guid>
    <title>JSON aus dem LLM erzwingen statt hoffen: Schema-Zwang</title>
    <link>https://www.ki-mittelstand.eu/blog/strukturierte-llm-ausgaben-json-schema</link>
    <description>Wer erzeugtes JSON parst und auf Gültigkeit hofft, baut eine Fehlerquelle ein. Wie Sie das Format technisch erzwingen — lokal und in der Cloud.</description>
    <pubDate>Tue, 04 Aug 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>json-schema</category><category>structured-output</category><category>vllm</category><category>automatisierung</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/vllm-cluster-multi-gpu-70b-modelle-self-hosted-2026</guid>
    <title>vLLM Cluster: 70B-Modelle auf Multi-GPU self-hosted</title>
    <link>https://www.ki-mittelstand.eu/blog/vllm-cluster-multi-gpu-70b-modelle-self-hosted-2026</link>
    <description>vLLM Cluster für Multi-GPU-Inferenz: Llama-3.3-70B auf 2-4 GPUs verteilen, Tensor-Parallelismus und Ray-Multi-Node konfigurieren — mit echten Configs.</description>
    <pubDate>Thu, 02 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>vllm</category><category>gpu-cluster</category><category>self-hosted</category><category>llm-inferenz</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/vllm-server-deutsch-anleitung-mittelstand</guid>
    <title>vLLM Server einrichten: Deutsch-Anleitung 2026</title>
    <link>https://www.ki-mittelstand.eu/blog/vllm-server-deutsch-anleitung-mittelstand</link>
    <description>vLLM auf eigenem Server installieren: GPU-Setup, API-Anbindung und Produktivbetrieb. Deutsche Anleitung für den Mittelstand.</description>
    <pubDate>Sun, 08 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>vllm</category><category>self-hosted-ki</category><category>mittelstand</category><category>deutschland</category><category>llm-server</category><category>on-premise</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/vllm-vs-ollama-durchsatz-benchmark-self-hosted-2026</guid>
    <title>vLLM vs Ollama: Durchsatz-Benchmark self-hosted 2026</title>
    <link>https://www.ki-mittelstand.eu/blog/vllm-vs-ollama-durchsatz-benchmark-self-hosted-2026</link>
    <description>vLLM vs Ollama im Durchsatz-Benchmark: bis 19x mehr Tokens/s unter Last durch Continuous Batching. Welches Tool wann self-hosted passt.</description>
    <pubDate>Thu, 02 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>vllm</category><category>ollama</category><category>self-hosted</category><category>llm-inference</category><category>benchmark</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/vllm-vs-sglang-vs-tensorrt-llm</guid>
    <title>vLLM vs. SGLang vs. TensorRT-LLM: welcher Server wofür</title>
    <link>https://www.ki-mittelstand.eu/blog/vllm-vs-sglang-vs-tensorrt-llm</link>
    <description>Drei Inferenzserver im Vergleich. Wo die Unterschiede real sind, wo sie in Benchmarks größer wirken als im Betrieb — und was Sie selbst messen müssen.</description>
    <pubDate>Tue, 04 Aug 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>vllm</category><category>sglang</category><category>tensorrt-llm</category><category>inferenz</category><category>durchsatz</category><category>mittelstand</category><category>deutschland</category>
  </item>

    </channel>
  </rss>
