
  <rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
      <title>KI im Mittelstand – News &amp; Technologie-Übersicht</title>
      <link>https://www.ki-mittelstand.eu/blog</link>
      <description>Faktenbasierte News, Technologien und Praxislösungen zu Künstlicher Intelligenz für mittelständische Unternehmen. Ohne Hype, mit klaren Quellen.</description>
      <language>de-de</language>
      <managingEditor>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</managingEditor>
      <webMaster>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</webMaster>
      <lastBuildDate>Sun, 17 May 2026 00:00:00 GMT</lastBuildDate>
      <atom:link href="https://www.ki-mittelstand.eu/tags/self-hosted/feed.xml" rel="self" type="application/rss+xml"/>
      
  <item>
    <guid>https://www.ki-mittelstand.eu/blog/azure-openai-vs-vllm-fuer-fertigung-500k-sparen-mit-tco-rech</guid>
    <title>Azure OpenAI vs. vLLM für Fertigung: €500k sparen mit TCO-Rechner 2026</title>
    <link>https://www.ki-mittelstand.eu/blog/azure-openai-vs-vllm-fuer-fertigung-500k-sparen-mit-tco-rech</link>
    <description>Azure OpenAI vs. vLLM im TCO-Vergleich: Wann sich Self-Hosting für die Fertigung lohnt und wie Sie den Break-Even konkret berechnen.</description>
    <pubDate>Sun, 17 May 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ki-technologie</category><category>mittelstand</category><category>deutschland</category><category>fertigung</category><category>qualitaetskontrolle</category><category>computer-vision</category><category>ausschuss</category><category>self-hosted</category><category>dsgvo</category><category>on-premise</category><category>azure-openai-tco</category><category>vllm-kosten-vergleich</category><category>cloud-vs-self-hosted-ai</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/chatbots-entwickeln-open-source-azure-openai-ai-search-deutschland-2025</guid>
    <title>RAG ChromaDB lokal: 85% weniger Halluzinationen</title>
    <link>https://www.ki-mittelstand.eu/blog/chatbots-entwickeln-open-source-azure-openai-ai-search-deutschland-2025</link>
    <description>RAG-Pipeline mit ChromaDB lokal: Firmenwissen durchsuchbar ohne Cloud. Unter €2.000 Hardware, 85% weniger Halluzinationen.</description>
    <pubDate>Mon, 09 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>rag</category><category>chromadb</category><category>wissensdatenbank</category><category>self-hosted</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/chatgpt-alternative-kostenlos-lokal-2026-kundenservice-prakt</guid>
    <title>ChatGPT-Alternative lokal: 5 Tools, €0 Abokosten</title>
    <link>https://www.ki-mittelstand.eu/blog/chatgpt-alternative-kostenlos-lokal-2026-kundenservice-prakt</link>
    <description>5 lokale ChatGPT-Alternativen ohne Abo: Ollama, LM Studio, GPT4All, Jan, LocalAI. 80-92% GPT-4-Qualität, €2.940/Jahr gespart bei 50 Nutzern.</description>
    <pubDate>Mon, 09 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>chatgpt-alternative</category><category>lokal</category><category>kostenlos</category><category>self-hosted</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/deepseek-coder-deutschland-praxisleitfaden-für-deutsche-kmu</guid>
    <title>DeepSeek Coder lokal: €12.000/Jahr vs. Copilot</title>
    <link>https://www.ki-mittelstand.eu/blog/deepseek-coder-deutschland-praxisleitfaden-für-deutsche-kmu</link>
    <description>DeepSeek Coder V2 lokal: €0 laufende Kosten statt €1.824/Jahr Copilot. DSGVO-konform, 16 GB VRAM, Setup in 2 Stunden.</description>
    <pubDate>Mon, 09 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>deepseek</category><category>coder</category><category>code-assistent</category><category>self-hosted</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/deepseek-offline-deutschland-praxisleitfaden-für-deutsche-km</guid>
    <title>Air-Gapped LLM: Llama 3.3 ohne Internet betreiben</title>
    <link>https://www.ki-mittelstand.eu/blog/deepseek-offline-deutschland-praxisleitfaden-für-deutsche-km</link>
    <description>Llama 3.3 komplett offline auf isoliertem Server: Setup in 4 Stunden, Hardware ab €3.200, €0 API-Kosten. Für KRITIS und Produktion.</description>
    <pubDate>Mon, 09 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>air-gapped</category><category>offline</category><category>llm</category><category>self-hosted</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/deepseek-v3-on-premise-chinas-open-source-llm</guid>
    <title>DeepSeek-v3 on-premise: Chinas Open-Source-LLM für deutsche Unternehmen</title>
    <link>https://www.ki-mittelstand.eu/blog/deepseek-v3-on-premise-chinas-open-source-llm</link>
    <description>DeepSeek-v3 auf eigenem Server betreiben: Chinas Open-Source-LLM mit 671B Parametern — für deutsche Unternehmen mit Datenschutz-Bedenken.</description>
    <pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>deepseek</category><category>open-weight</category><category>chinese-llm</category><category>self-hosted</category><category>ki-server</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/deepseek-vs-chatgpt-deutschland-praxisleitfaden-für-deutsche</guid>
    <title>DeepSeek R1 lokal: 90% GPT-4-Qualität, DSGVO-ok</title>
    <link>https://www.ki-mittelstand.eu/blog/deepseek-vs-chatgpt-deutschland-praxisleitfaden-für-deutsche</link>
    <description>DeepSeek R1 in 45 Min. lokal installieren: 85-90% GPT-4-Qualität, Hardware ab €3.500, keine Cloud-Kosten. DSGVO-konform.</description>
    <pubDate>Mon, 09 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>deepseek</category><category>lokal</category><category>dsgvo</category><category>self-hosted</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/eigener-ki-server-hardware-guide-2026-gpu-server-für-deutsch</guid>
    <title>KI-Server kaufen 2026: 3 Konfigurationen ab €3.500</title>
    <link>https://www.ki-mittelstand.eu/blog/eigener-ki-server-hardware-guide-2026-gpu-server-für-deutsch</link>
    <description>KI-Server kaufen für RAG, Chatbots, Code-Assistenten: 3 GPU-Konfigurationen ab €3.500 mit Benchmarks (RTX 4090, A4000, L4) und Einkaufsliste.</description>
    <pubDate>Mon, 09 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ki-server</category><category>gpu</category><category>hardware</category><category>self-hosted</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/gemma-2-27b-eigenen-server-google-open-weight</guid>
    <title>Gemma 2 27B auf eigenem Server: Googles Open-Weight-Modell on-premise</title>
    <link>https://www.ki-mittelstand.eu/blog/gemma-2-27b-eigenen-server-google-open-weight</link>
    <description>Gemma 2 27B von Google auf dem eigenen Server: Open-Weight-Modell mit hervorragender Sprachqualität. Hardware, Performance und Deployment-Guide.</description>
    <pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>gemma</category><category>google</category><category>open-weight</category><category>self-hosted</category><category>ki-server</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/github-copilot-alternative-lokal-fuer-fertigung-150k-ausschu</guid>
    <title>GitHub Copilot Alternative: Code-KI lokal hosten</title>
    <link>https://www.ki-mittelstand.eu/blog/github-copilot-alternative-lokal-fuer-fertigung-150k-ausschu</link>
    <description>GitHub Copilot Alternative lokal hosten: DSGVO-konforme KI-Code-Assistenz auf eigener Hardware mit Open-Source-LLMs. Code bleibt im Unternehmen.</description>
    <pubDate>Thu, 21 May 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ki-technologie</category><category>mittelstand</category><category>deutschland</category><category>fertigung</category><category>qualitaetskontrolle</category><category>computer-vision</category><category>ausschuss</category><category>self-hosted</category><category>dsgvo</category><category>on-premise</category><category>copilot-alternative-lokal</category><category>private-code-ai</category><category>code-assist-dsgvo</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/github-copilot-alternative-lokal-fuer-fertigung-200k-einspar</guid>
    <title>Lokale Copilot-Alternative: private Code-KI 2026</title>
    <link>https://www.ki-mittelstand.eu/blog/github-copilot-alternative-lokal-fuer-fertigung-200k-einspar</link>
    <description>GitHub Copilot Alternative lokal für die Fertigung. Senken Sie mit Private Code AI Ihre Entwicklungskosten um €200.000 und schützen Sie geistiges Eigentum.</description>
    <pubDate>Thu, 21 May 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ki-technologie</category><category>mittelstand</category><category>deutschland</category><category>fertigung</category><category>qualitaetskontrolle</category><category>computer-vision</category><category>ausschuss</category><category>self-hosted</category><category>dsgvo</category><category>on-premise</category><category>copilot-alternative-lokal</category><category>private-code-ai</category><category>code-assist-dsgvo</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/glm-4-eigenen-server-zhipuai-open-source-llm</guid>
    <title>GLM-4 auf eigenem Server: ZhipuAI Open-Source-LLM für deutsche Unternehmen</title>
    <link>https://www.ki-mittelstand.eu/blog/glm-4-eigenen-server-zhipuai-open-source-llm</link>
    <description>GLM-4 von ZhipuAI auf eigenem Server betreiben: Das Open-Source-LLM der Tsinghua-Universität mit 128K-Kontext für deutsche Unternehmen — on-premise und DSGVO-konform.</description>
    <pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>glm</category><category>zhipuai</category><category>open-weight</category><category>self-hosted</category><category>ki-server</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/hashicorp-vault-azure-hochverfuegbare-secrets-fuer-fertigung</guid>
    <title>HashiCorp Vault auf Azure: Secrets hochverfügbar</title>
    <link>https://www.ki-mittelstand.eu/blog/hashicorp-vault-azure-hochverfuegbare-secrets-fuer-fertigung</link>
    <description>HashiCorp Vault auf Azure hochverfügbar betreiben: Secrets-Management mit Zero-Trust, HA-Cluster und reduziertem Ausfallrisiko für kritische Systeme.</description>
    <pubDate>Sun, 07 Jun 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ki-technologie</category><category>mittelstand</category><category>deutschland</category><category>fertigung</category><category>qualitaetskontrolle</category><category>computer-vision</category><category>ausschuss</category><category>self-hosted</category><category>dsgvo</category><category>on-premise</category><category>hashicorp-vault-azure</category><category>secrets-management-enterprise</category><category>zero-trust-vault</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/hetzner-gpu-cloud-ki-server-mieten-statt-kaufen</guid>
    <title>Hetzner GPU-Cloud: KI-Server mieten statt kaufen</title>
    <link>https://www.ki-mittelstand.eu/blog/hetzner-gpu-cloud-ki-server-mieten-statt-kaufen</link>
    <description>Hetzner GPU-Cloud: KI-Workloads auf Dedicated-Server-Rental. RTX 4090, A100 und H100-Instanzen — wann sich Cloud-Lösung für deutsche Unternehmen lohnt.</description>
    <pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>hetzner</category><category>gpu-cloud</category><category>ki-server-mieten</category><category>self-hosted</category><category>mittelstand</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/hetzner-gpu-server-ki-workloads-cloud</guid>
    <title>Hetzner GPU-Server: KI-Workloads mit der Cloud</title>
    <link>https://www.ki-mittelstand.eu/blog/hetzner-gpu-server-ki-workloads-cloud</link>
    <description>Hetzner GPU-Server für KI-Workloads: RTX 4090, A4000 und H100-Instanzen gemietet statt gekauft. Wann sich Cloud lohnt, wann der On-Premise-Server günstiger ist.</description>
    <pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>hetzner</category><category>gpu-server</category><category>cloud</category><category>self-hosted</category><category>ki-server</category><category>mittelstand</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/iso-27001-ki-audit-fertigung-250k-einsparung-mit-checkliste</guid>
    <title>ISO 27001 KI-Audit: Checkliste zur Zertifizierung</title>
    <link>https://www.ki-mittelstand.eu/blog/iso-27001-ki-audit-fertigung-250k-einsparung-mit-checkliste</link>
    <description>ISO 27001 KI-Audit für Self-Hosted-AI: Checkliste zu Pod Security, mTLS, SBOMs und Gap-Analyse. So bereiten Sie die Zertifizierung strukturiert vor.</description>
    <pubDate>Fri, 29 May 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ki-technologie</category><category>mittelstand</category><category>deutschland</category><category>fertigung</category><category>qualitaetskontrolle</category><category>computer-vision</category><category>ausschuss</category><category>self-hosted</category><category>dsgvo</category><category>on-premise</category><category>eu-ai-act</category><category>iso-27001-ki</category><category>ai-audit-vorbereitung</category><category>self-hosted-ai-zertifizierung</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/ki-energieversorgungsunternehmen-deutschland-praktischer-lei</guid>
    <title>OpenWebUI Benutzergruppen: LDAP + Rollen</title>
    <link>https://www.ki-mittelstand.eu/blog/ki-energieversorgungsunternehmen-deutschland-praktischer-lei</link>
    <description>OpenWebUI Benutzergruppen mit LDAP und Rollensteuerung einrichten. Modellzugriff pro Team steuern, Token-Limits setzen — in 2-3 Stunden.</description>
    <pubDate>Sun, 08 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>openwebui</category><category>benutzerverwaltung</category><category>self-hosted</category><category>chat</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/ki-kosten-fertigung-cloud-ausgaben-senken-von-4800-auf-400-l</guid>
    <title>KI Kosten Fertigung: Cloud-Ausgaben senken von €4.800 auf €400 lokal 2026</title>
    <link>https://www.ki-mittelstand.eu/blog/ki-kosten-fertigung-cloud-ausgaben-senken-von-4800-auf-400-l</link>
    <description>Fertigungsunternehmen können KI-Kosten drastisch senken, indem sie von Cloud-Lösungen auf On-Premise-Systeme umsteigen und so monatlich bis zu 92% sparen. Ein Praxisleitfaden mit ROI-Berechnung für 2026.</description>
    <pubDate>Wed, 25 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ki-technologie</category><category>mittelstand</category><category>deutschland</category><category>fertigung</category><category>qualitaetskontrolle</category><category>computer-vision</category><category>ausschuss</category><category>self-hosted</category><category>dsgvo</category><category>on-premise</category><category>ki-kosten-vergleich</category><category>cloud-vs-on-premise-kosten</category><category>ai-tco-rechner</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/ki-migration-fertigung-300000-kostenersparnis-durch-azure-op</guid>
    <title>KI-Migration: von Azure OpenAI zu Self-Hosted</title>
    <link>https://www.ki-mittelstand.eu/blog/ki-migration-fertigung-300000-kostenersparnis-durch-azure-op</link>
    <description>Sparen Sie als Fertigungsunternehmen bis zu 300.000 € pro Jahr durch die Migration von Azure OpenAI zu einer Self-Hosted-Lösung. Unser Playbook für den reibungslosen Übergang in 30 Tagen.</description>
    <pubDate>Tue, 07 Apr 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ki-technologie</category><category>mittelstand</category><category>deutschland</category><category>fertigung</category><category>qualitaetskontrolle</category><category>computer-vision</category><category>ausschuss</category><category>self-hosted</category><category>dsgvo</category><category>on-premise</category><category>azure-zu-self-hosted-migration</category><category>openai-alternative-migration</category><category>cloud-exit-ai</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/ki-ohne-cloud-on-premise-deutschland-2026-self-hosted-prakti</guid>
    <title>Self-Hosted LLM: Kosten-Guide für Ollama &amp; vLLM 2026</title>
    <link>https://www.ki-mittelstand.eu/blog/ki-ohne-cloud-on-premise-deutschland-2026-self-hosted-prakti</link>
    <description>Vergleichen Sie die Self-Hosting-Kosten von Open-Source-LLMs. Unser Praxis-Guide zeigt TCO für Ollama vs. vLLM für 100+ Mitarbeiter ab 500€/Monat.</description>
    <pubDate>Mon, 09 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>on-premise</category><category>self-hosted</category><category>llm</category><category>vergleich</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/ki-on-premise-einstieg-mittelstand</guid>
    <title>KI on-premise betreiben: der Einstieg</title>
    <link>https://www.ki-mittelstand.eu/blog/ki-on-premise-einstieg-mittelstand</link>
    <description>Was KI on-premise wirklich verlangt: Hardware, Stellplatz, Betrieb, Update-Konzept — und die drei Einstiegsszenarien, die in vier Wochen stehen.</description>
    <pubDate>Sun, 02 Aug 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>on-premise</category><category>ki-server</category><category>self-hosted</category><category>grundlagen</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/ki-server-docker-containerisierung-llm-pipelines</guid>
    <title>KI-Server-Docker: Containerisierung für LLM-Pipelines</title>
    <link>https://www.ki-mittelstand.eu/blog/ki-server-docker-containerisierung-llm-pipelines</link>
    <description>KI-Server mit Docker und Docker Compose: LLM, Vektordatenbank und Frontend in Containern — reproduzierbar, skalierbar und einfach zu warten.</description>
    <pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>docker</category><category>container</category><category>ki-server</category><category>devops</category><category>self-hosted</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/ki-server-kosten-laufende-rechnung</guid>
    <title>KI-Server Kosten: die Rechnung nach dem Kauf</title>
    <link>https://www.ki-mittelstand.eu/blog/ki-server-kosten-laufende-rechnung</link>
    <description>Strom, Kühlung, Wartung, Ersatzteile: was ein KI-Server über 36 Monate wirklich kostet — mit offengelegter Rechnung und Kipp-Punkt gegen Cloud-APIs.</description>
    <pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ki-server</category><category>kosten</category><category>tco</category><category>self-hosted</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/ki-server-rag-vom-pdf-zur-wissensdatenbank</guid>
    <title>KI-Server-RAG: Vom PDF zur abfragebereiten Wissensdatenbank</title>
    <link>https://www.ki-mittelstand.eu/blog/ki-server-rag-vom-pdf-zur-wissensdatenbank</link>
    <description>RAG-Pipeline von PDF zur Wissensdatenbank: Dokumenten-Extraktion, Chunking, Embedding und Vektorsuche — komplette Implementierung.</description>
    <pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>rag</category><category>pdf</category><category>vektorsuche</category><category>mittelstand</category><category>self-hosted</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/ki-server-selbst-bauen-gpu-workstation-llm-2026</guid>
    <title>KI-Server selbst bauen: GPU-Workstation für LLMs 2026</title>
    <link>https://www.ki-mittelstand.eu/blog/ki-server-selbst-bauen-gpu-workstation-llm-2026</link>
    <description>GPU-Workstation für LLMs selbst bauen: GPU, CPU, RAM und Netzteil mit realen Preisen ab €1.800 — 3 Budget-Konfigurationen für 7B bis 70B.</description>
    <pubDate>Thu, 02 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ki-server</category><category>gpu</category><category>hardware</category><category>self-hosted</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/ki-spracherkennung-und-uebersetzung</guid>
    <title>Whisper lokal: 95 % Genauigkeit ohne Cloud</title>
    <link>https://www.ki-mittelstand.eu/blog/ki-spracherkennung-und-uebersetzung</link>
    <description>Whisper lokal installieren und deutsche Meetings DSGVO-konform transkribieren. 95 % Genauigkeit, einmalig €800-1.500 statt Cloud-Abo.</description>
    <pubDate>Mon, 09 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>whisper</category><category>transkription</category><category>self-hosted</category><category>meetings</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/llama3.3-70b-eigenen-server-complete-guide</guid>
    <title>Llama3.3 70B auf eigenem Server: Der Complete Guide 2026</title>
    <link>https://www.ki-mittelstand.eu/blog/llama3.3-70b-eigenen-server-complete-guide</link>
    <description>Llama3.3 70B von Meta auf dem eigenen Server betreiben — Hardware, Kosten, Performance und Deployment-Guide für den deutschen Mittelstand.</description>
    <pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>llama</category><category>meta</category><category>open-weight</category><category>self-hosted</category><category>ki-server</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/llmlite-vs-ollama-lokale-enterprise-ki</guid>
    <title>LiteLLM Proxy: 30-40 % API-Kosten sparen</title>
    <link>https://www.ki-mittelstand.eu/blog/llmlite-vs-ollama-lokale-enterprise-ki</link>
    <description>LiteLLM Proxy bündelt OpenAI, Anthropic und Ollama unter einer API. Setup in unter 1 Stunde, 30-40 % weniger API-Kosten durch Routing.</description>
    <pubDate>Mon, 09 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>litellm</category><category>proxy</category><category>llm</category><category>self-hosted</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/localai-fuer-fertigung-120000-einsparung-durch-eigene-openai</guid>
    <title>LocalAI: eigene OpenAI-kompatible API hosten</title>
    <link>https://www.ki-mittelstand.eu/blog/localai-fuer-fertigung-120000-einsparung-durch-eigene-openai</link>
    <description>Mit LocalAI hosten Sie eine OpenAI-kompatible API selbst, DSGVO-konform und on-premise. Installation per Docker, kompatibel mit gängigen LLMs und SDKs.</description>
    <pubDate>Fri, 27 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ki-technologie</category><category>mittelstand</category><category>deutschland</category><category>fertigung</category><category>qualitaetskontrolle</category><category>computer-vision</category><category>ausschuss</category><category>self-hosted</category><category>dsgvo</category><category>on-premise</category><category>localai-installation</category><category>openai-api-alternative</category><category>self-hosted-gpt-api</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/localai-production-fuer-fertigung-500k-einsparung-durch-open</guid>
    <title>LocalAI in Produktion: OpenAI lokal ersetzen</title>
    <link>https://www.ki-mittelstand.eu/blog/localai-production-fuer-fertigung-500k-einsparung-durch-open</link>
    <description>LocalAI in Produktion: OpenAI durch eine API-kompatible Open-Source-Lösung ersetzen. Drop-in-Migration, sensible Daten bleiben lokal und DSGVO-konform.</description>
    <pubDate>Mon, 29 Jun 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ki-technologie</category><category>mittelstand</category><category>deutschland</category><category>fertigung</category><category>qualitaetskontrolle</category><category>computer-vision</category><category>ausschuss</category><category>self-hosted</category><category>dsgvo</category><category>on-premise</category><category>localai-enterprise</category><category>openai-alternative</category><category>self-hosted-gpt-api</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/mamba-hybrid-kv-cache-langer-kontext</guid>
    <title>Mamba-Hybrid: KV-Cache bei langem Kontext</title>
    <link>https://www.ki-mittelstand.eu/blog/mamba-hybrid-kv-cache-langer-kontext</link>
    <description>Der KV-Cache frisst bei 32k Kontext mehr VRAM als das Modell selbst. Was Mamba-Hybride wie Soofi S daran ändern — und wann FP8 die billigere Antwort ist.</description>
    <pubDate>Sun, 02 Aug 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>llm-inferenz</category><category>vram</category><category>self-hosted</category><category>rag</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/microsoft-copilot-kosten-sharepoint-anbindung</guid>
    <title>Copilot vs. lokale KI: TCO bei 100 Nutzern</title>
    <link>https://www.ki-mittelstand.eu/blog/microsoft-copilot-kosten-sharepoint-anbindung</link>
    <description>Microsoft Copilot kostet €42.000/Jahr bei 100 Nutzern. Lokale KI-Alternative: €18.000/Jahr mit voller DSGVO-Konformität. TCO-Rechnung.</description>
    <pubDate>Mon, 09 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>copilot</category><category>tco</category><category>vergleich</category><category>self-hosted</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/mistral-large-2-on-premise-französische-ki-deutschland</guid>
    <title>Mistral Large 2 on-premise: Französische KI auf deutschem Server</title>
    <link>https://www.ki-mittelstand.eu/blog/mistral-large-2-on-premise-französische-ki-deutschland</link>
    <description>Mistral Large 2 von der französischen Startup auf dem eigenen Server — europäische Alternative zu US-Modellen, DSGVO-konform und on-premise.</description>
    <pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>mistral</category><category>frankreich</category><category>open-weight</category><category>self-hosted</category><category>ki-server</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/mlflow-kubernetes-experiment-tracking-fuer-fertigung-450k-ei</guid>
    <title>MLflow auf Kubernetes: Experiment-Tracking self-hosted</title>
    <link>https://www.ki-mittelstand.eu/blog/mlflow-kubernetes-experiment-tracking-fuer-fertigung-450k-ei</link>
    <description>MLflow auf Kubernetes betreiben: ML-Experimente zentral tracken, Modelle versionieren und die Registry self-hosted absichern. Setup mit echten Manifesten.</description>
    <pubDate>Thu, 02 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ki-technologie</category><category>mittelstand</category><category>deutschland</category><category>fertigung</category><category>qualitaetskontrolle</category><category>computer-vision</category><category>ausschuss</category><category>self-hosted</category><category>dsgvo</category><category>on-premise</category><category>mlflow-enterprise</category><category>kubernetes-mlops</category><category>experiment-tracking</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/ollama-cluster-mehrere-rechner-sharding</guid>
    <title>Ollama-Cluster über mehrere Rechner: Sharding</title>
    <link>https://www.ki-mittelstand.eu/blog/ollama-cluster-mehrere-rechner-sharding</link>
    <description>Ein LLM über zwei Rechner verteilen: Ollama kann es nicht, llama.cpp schon. Befehle, echte Benchmarks — und warum meist eine größere GPU gewinnt.</description>
    <pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ollama</category><category>llama-cpp</category><category>self-hosted</category><category>gpu</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/ollama-kubernetes-deployen-llm-cluster-2026</guid>
    <title>Ollama auf Kubernetes: LLM-Cluster mit Autoscaling</title>
    <link>https://www.ki-mittelstand.eu/blog/ollama-kubernetes-deployen-llm-cluster-2026</link>
    <description>Ollama auf Kubernetes deployen: StatefulSet, NVIDIA GPU-Scheduling, Helm Chart und HPA-Autoscaling für einen produktiven LLM-Cluster.</description>
    <pubDate>Thu, 02 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ollama</category><category>kubernetes</category><category>gpu</category><category>self-hosted</category><category>infrastruktur</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/ollama-load-balancer-ohne-kubernetes</guid>
    <title>Ollama Load-Balancer ohne Kubernetes einrichten</title>
    <link>https://www.ki-mittelstand.eu/blog/ollama-load-balancer-ohne-kubernetes</link>
    <description>Ollama serialisiert Anfragen pro Modell. Mehrere Instanzen hinter nginx mit least_conn: systemd-Units, Config und die Timeouts, die wirklich zählen.</description>
    <pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ollama</category><category>self-hosted</category><category>gpu</category><category>nginx</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/ollama-synology-nas-installieren</guid>
    <title>Ollama auf Synology NAS: Docker Setup in 30 Min</title>
    <link>https://www.ki-mittelstand.eu/blog/ollama-synology-nas-installieren</link>
    <description>Ollama auf Synology NAS via Docker installieren in 30 Minuten. Phi-3 antwortet in 3-8 Sek., 7B-Modelle ab 16 GB RAM. Ohne Cloud.</description>
    <pubDate>Sun, 08 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ollama</category><category>synology</category><category>nas</category><category>docker</category><category>self-hosted</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/ollama-ubuntu-installieren-anleitung-2025-self-hosted-prakti</guid>
    <title>Ollama Ubuntu installieren: LLM lokal 15 Min</title>
    <link>https://www.ki-mittelstand.eu/blog/ollama-ubuntu-installieren-anleitung-2025-self-hosted-prakti</link>
    <description>Ollama auf Ubuntu installieren: Lokales LLM in 15 Minuten. Llama 3.1 auf eigenem Server, €0 API-Kosten, volle DSGVO-Kontrolle.</description>
    <pubDate>Mon, 09 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ollama</category><category>ubuntu</category><category>self-hosted</category><category>llm</category><category>installation</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/ollama-ubuntu-installieren-anleitung-2026-self-hosted-prakti</guid>
    <title>Ollama Cluster: Load Balancing für 200+ Nutzer</title>
    <link>https://www.ki-mittelstand.eu/blog/ollama-ubuntu-installieren-anleitung-2026-self-hosted-prakti</link>
    <description>Ollama Cluster mit Load Balancing: 200+ Nutzer, automatisches Failover, horizontale Skalierung. Nginx-Setup für den Mittelstand.</description>
    <pubDate>Mon, 09 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ollama</category><category>cluster</category><category>load-balancing</category><category>self-hosted</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/ollama-ubuntu-private-ki-konfiguration</guid>
    <title>Ollama GPU CUDA Setup: Ubuntu Server Anleitung</title>
    <link>https://www.ki-mittelstand.eu/blog/ollama-ubuntu-private-ki-konfiguration</link>
    <description>Ollama mit NVIDIA GPU und CUDA auf Ubuntu: 8x schneller als CPU. Anleitung für CUDA-Treiber, VRAM-Optimierung und Produktion.</description>
    <pubDate>Mon, 09 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ollama</category><category>gpu</category><category>cuda</category><category>ubuntu</category><category>self-hosted</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/ollama-vs-vllm-vs-lm-studio-llm-frontendvergleich</guid>
    <title>Ollama vs vLLM vs LM Studio: Welches LLM-Frontend für den Mittelstand?</title>
    <link>https://www.ki-mittelstand.eu/blog/ollama-vs-vllm-vs-lm-studio-llm-frontendvergleich</link>
    <description>Ollama, vLLM und LM Studio im Vergleich: Welches LLM-Frontend passt für den deutschen Mittelstand? Performance, Features, Deployment und Kosten.</description>
    <pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ollama</category><category>vllm</category><category>lm-studio</category><category>llm-frontends</category><category>self-hosted</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/ollama-vs-vllm-vs-localai-llm-server-fuer-die-fertigung-250k</guid>
    <title>Ollama vs vLLM vs LocalAI: LLM-Server für die Fertigung – €250k Kosten sparen 2026</title>
    <link>https://www.ki-mittelstand.eu/blog/ollama-vs-vllm-vs-localai-llm-server-fuer-die-fertigung-250k</link>
    <description>Ollama vs. vLLM vs. LocalAI: LLM-Server im Benchmark zu Durchsatz, GPU-Effizienz und API-Kompatibilität. Welcher Self-Hosted-Server wann passt.</description>
    <pubDate>Sun, 14 Jun 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ki-technologie</category><category>mittelstand</category><category>deutschland</category><category>fertigung</category><category>qualitaetskontrolle</category><category>computer-vision</category><category>ausschuss</category><category>self-hosted</category><category>dsgvo</category><category>on-premise</category><category>ollama-vllm-vergleich</category><category>llm-server-benchmark</category><category>localai-alternative</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/openwebui-installieren-docker-anleitung-2026-self-hosted-pra</guid>
    <title>OpenWebUI Teams: Rollen und API-Keys verwalten</title>
    <link>https://www.ki-mittelstand.eu/blog/openwebui-installieren-docker-anleitung-2026-self-hosted-pra</link>
    <description>OpenWebUI für Teams: Rollen, API-Keys und Berechtigungen verwalten. Mit LDAP-Anbindung und Kostencontrolling pro Nutzer.</description>
    <pubDate>Mon, 09 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>openwebui</category><category>teams</category><category>rollen</category><category>self-hosted</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/openwebui-private-ki-chatbots-im-unternehmen</guid>
    <title>OpenWebUI + Ollama: Firmen-ChatGPT in 30 Min</title>
    <link>https://www.ki-mittelstand.eu/blog/openwebui-private-ki-chatbots-im-unternehmen</link>
    <description>OpenWebUI und Ollama als Firmen-ChatGPT: Multi-User, RAG und DSGVO-konform für €89/Monat. Docker-Compose-Anleitung.</description>
    <pubDate>Mon, 09 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>openwebui</category><category>ollama</category><category>firmen-chatgpt</category><category>self-hosted</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/pinecone-zu-qdrant-migration-150k-kosten-sparen-im-fertigung</guid>
    <title>Pinecone zu Qdrant Migration: €150k Kosten sparen im Fertigungs-Mittelstand 2026</title>
    <link>https://www.ki-mittelstand.eu/blog/pinecone-zu-qdrant-migration-150k-kosten-sparen-im-fertigung</link>
    <description>Pinecone zu Qdrant migrieren: Vektor-Datenbank self-hosten per API-Export und DNS-Umstellung, Daten bleiben in der EU. Praxisleitfaden für Fertiger.</description>
    <pubDate>Tue, 19 May 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ki-technologie</category><category>mittelstand</category><category>deutschland</category><category>fertigung</category><category>qualitaetskontrolle</category><category>computer-vision</category><category>ausschuss</category><category>self-hosted</category><category>dsgvo</category><category>on-premise</category><category>pinecone-alternative</category><category>qdrant-migration</category><category>vector-db-self-hosted</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/qdrant-cluster-fertigung-350k-ausschuss-durch-ki-vektorsuche</guid>
    <title>Qdrant-Cluster aufsetzen: skalierbare Vektorsuche</title>
    <link>https://www.ki-mittelstand.eu/blog/qdrant-cluster-fertigung-350k-ausschuss-durch-ki-vektorsuche</link>
    <description>Qdrant-Cluster on-premise aufsetzen: hochverfügbare Vektorsuche für RAG und Bildklassifizierung. Architektur, Replikation und Betrieb Schritt für Schritt.</description>
    <pubDate>Sun, 28 Jun 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ki-technologie</category><category>mittelstand</category><category>deutschland</category><category>fertigung</category><category>qualitaetskontrolle</category><category>computer-vision</category><category>ausschuss</category><category>self-hosted</category><category>dsgvo</category><category>on-premise</category><category>qdrant-cluster</category><category>vektor-db-self-hosted</category><category>high-availability-vector-search</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/qwen-eigenen-server-chinas-open-weight-modell-on-premise</guid>
    <title>Qwen auf eigenem Server: Chinas Open-Weight-Modell on-premise betreiben</title>
    <link>https://www.ki-mittelstand.eu/blog/qwen-eigenen-server-chinas-open-weight-modell-on-premise</link>
    <description>Qwen von Alibaba als Open-Weight-Modell auf dem eigenen KI-Server betreiben: Warum deutsche Unternehmen jetzt auf chinesische LLMs setzen – DSGVO-konform, komplett on-premise.</description>
    <pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>qwen</category><category>open-weight</category><category>llm</category><category>self-hosted</category><category>ki-server</category><category>mittelstand</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/rag-mit-langchain-python-guide-deutscher-mittelstand</guid>
    <title>RAG mit LangChain: Python-Guide für den deutschen Mittelstand</title>
    <link>https://www.ki-mittelstand.eu/blog/rag-mit-langchain-python-guide-deutscher-mittelstand</link>
    <description>RAG mit LangChain in Python: Von PDF zu abfragebereitem Wissensbestand — vollständiger Guide für den deutschen Mittelstand mit Open-Source-LLMs.</description>
    <pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>rag</category><category>langchain</category><category>python</category><category>mittelstand</category><category>self-hosted</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/self-hosted-vs-cloud-ki-vergleich</guid>
    <title>Self-Hosted vs Cloud KI: Der ehrliche Vergleich 2026</title>
    <link>https://www.ki-mittelstand.eu/blog/self-hosted-vs-cloud-ki-vergleich</link>
    <description>Self-Hosted vs Cloud KI: Der ehrliche Vergleich für deutsche Unternehmen — Kosten, Datenschutz, Performance und wann welche Lösung passt.</description>
    <pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>self-hosted</category><category>cloud</category><category>vergleich</category><category>ki</category><category>mittelstand</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/server-fuer-ki-anwendungen-sizing</guid>
    <title>Server für KI-Anwendungen: Sizing nach Use-Case</title>
    <link>https://www.ki-mittelstand.eu/blog/server-fuer-ki-anwendungen-sizing</link>
    <description>RAG, Chat, OCR, Code oder Sichtprüfung — jede Anwendung braucht anderes Sizing. Fünf Faustformeln plus die KV-Cache-Rechnung, die 70B-Pläne killt.</description>
    <pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ki-server</category><category>hardware</category><category>rag</category><category>self-hosted</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/soofi-s-selbst-hosten-gpu</guid>
    <title>Soofi S selbst hosten: welche GPU reicht wirklich?</title>
    <link>https://www.ki-mittelstand.eu/blog/soofi-s-selbst-hosten-gpu</link>
    <description>Soofi S hat 3,2 Mrd. aktive Parameter, braucht aber VRAM für 31,6 Mrd. Was das für die GPU-Wahl heißt — mit offener Rechnung und Stand der Tools.</description>
    <pubDate>Sun, 02 Aug 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>soofi</category><category>llm</category><category>gpu</category><category>self-hosted</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/soofi-s-vs-llama-qwen-deutsch</guid>
    <title>Soofi S vs Llama 3 vs Qwen für deutschen Text</title>
    <link>https://www.ki-mittelstand.eu/blog/soofi-s-vs-llama-qwen-deutsch</link>
    <description>Soofi S, Llama 3.3 und Qwen3.5 im Vergleich für deutsche Texte: Lizenz, VRAM, Verfügbarkeit — und ein Testaufbau mit den eigenen Dokumenten.</description>
    <pubDate>Sun, 02 Aug 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>llm</category><category>open-source</category><category>self-hosted</category><category>soofi</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/souveraene-ki-definition-mittelstand</guid>
    <title>Souveräne KI: was der Begriff technisch heißt</title>
    <link>https://www.ki-mittelstand.eu/blog/souveraene-ki-definition-mittelstand</link>
    <description>Datenresidenz, Betriebssouveränität, Technologiewahl: was souveräne KI wirklich meint — mit Prüfliste für Anbieter und offener Kostenrechnung.</description>
    <pubDate>Sun, 02 Aug 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>souveraenitaet</category><category>compliance</category><category>cloud</category><category>self-hosted</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/tabby-ml-fuer-fertigung-code-assistent-fuer-70000-ersparnis</guid>
    <title>Tabby ML: Code-Vervollständigung für Entwicklerteams</title>
    <link>https://www.ki-mittelstand.eu/blog/tabby-ml-fuer-fertigung-code-assistent-fuer-70000-ersparnis</link>
    <description>Ersparen Sie der Fertigung bis zu 70.000 € pro Jahr mit einem lokal gehosteten KI-Code-Assistenten wie Tabby ML. Maximale Code-Privatsphäre und Kontrolle für deutsche Mittelständler.</description>
    <pubDate>Sat, 28 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ki-technologie</category><category>mittelstand</category><category>deutschland</category><category>fertigung</category><category>qualitaetskontrolle</category><category>computer-vision</category><category>ausschuss</category><category>self-hosted</category><category>dsgvo</category><category>on-premise</category><category>tabby-ml-installation</category><category>github-copilot-alternative</category><category>code-ai-lokal</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/vllm-cluster-multi-gpu-70b-modelle-self-hosted-2026</guid>
    <title>vLLM Cluster: 70B-Modelle auf Multi-GPU self-hosted</title>
    <link>https://www.ki-mittelstand.eu/blog/vllm-cluster-multi-gpu-70b-modelle-self-hosted-2026</link>
    <description>vLLM Cluster für Multi-GPU-Inferenz: Llama-3.3-70B auf 2-4 GPUs verteilen, Tensor-Parallelismus und Ray-Multi-Node konfigurieren — mit echten Configs.</description>
    <pubDate>Thu, 02 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>vllm</category><category>gpu-cluster</category><category>self-hosted</category><category>llm-inferenz</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/vllm-server-enterprise-setup-2025-gpu-server-für-deutsche-un</guid>
    <title>Whisper API vs lokal: Kosten pro Audiostunde</title>
    <link>https://www.ki-mittelstand.eu/blog/vllm-server-enterprise-setup-2025-gpu-server-für-deutsche-un</link>
    <description>Whisper API vs. Self-Hosted: Ab 80 Audiostunden/Monat lohnt der eigene Server – €0,02 statt €0,36 pro Minute Transkription.</description>
    <pubDate>Mon, 09 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>whisper</category><category>spracherkennung</category><category>transkription</category><category>kosten</category><category>self-hosted</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/vllm-server-enterprise-setup-2025-praktischer-leitfaden-für</guid>
    <title>LocalAI auf Raspberry Pi: KI für €80 Hardware</title>
    <link>https://www.ki-mittelstand.eu/blog/vllm-server-enterprise-setup-2025-praktischer-leitfaden-für</link>
    <description>LocalAI auf Raspberry Pi 5: Kleine KI-Modelle lokal ausführen für €80 Hardware. Setup für Textklassifikation und Embeddings.</description>
    <pubDate>Mon, 09 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>localai</category><category>raspberry-pi</category><category>edge-ai</category><category>inferenz</category><category>self-hosted</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/vllm-vs-ollama-durchsatz-benchmark-self-hosted-2026</guid>
    <title>vLLM vs Ollama: Durchsatz-Benchmark self-hosted 2026</title>
    <link>https://www.ki-mittelstand.eu/blog/vllm-vs-ollama-durchsatz-benchmark-self-hosted-2026</link>
    <description>vLLM vs Ollama im Durchsatz-Benchmark: bis 19x mehr Tokens/s unter Last durch Continuous Batching. Welches Tool wann self-hosted passt.</description>
    <pubDate>Thu, 02 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>vllm</category><category>ollama</category><category>self-hosted</category><category>llm-inference</category><category>benchmark</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/was-ist-ein-ki-server-definition-mittelstand</guid>
    <title>Was ist ein KI-Server? Definition und Einsatz</title>
    <link>https://www.ki-mittelstand.eu/blog/was-ist-ein-ki-server-definition-mittelstand</link>
    <description>KI-Server erklärt: Was ihn von einem normalen Server unterscheidet, welche drei Bauformen es gibt und wann die Cloud die bessere Wahl bleibt.</description>
    <pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ki-server</category><category>gpu</category><category>self-hosted</category><category>grundlagen</category><category>mittelstand</category><category>deutschland</category>
  </item>

  <item>
    <guid>https://www.ki-mittelstand.eu/blog/whisper-lokal-fuer-fertigung-70000-einsparung-durch-ki-sprac</guid>
    <title>Whisper lokal: KI-Spracherkennung ohne Cloud</title>
    <link>https://www.ki-mittelstand.eu/blog/whisper-lokal-fuer-fertigung-70000-einsparung-durch-ki-sprac</link>
    <description>Whisper lokal einrichten: DSGVO-konforme KI-Spracherkennung ohne Cloud-Kosten. Berichte und Maschinendaten präzise transkribieren und automatisch auswerten.</description>
    <pubDate>Tue, 24 Mar 2026 00:00:00 GMT</pubDate>
    <author>phillip.pham@pexon-consulting.de (KI Mittelstand Team)</author>
    <category>ki-technologie</category><category>mittelstand</category><category>deutschland</category><category>fertigung</category><category>qualitaetskontrolle</category><category>spracherkennung-lokal</category><category>self-hosted</category><category>dsgvo</category><category>on-premise</category><category>whisper-self-hosted</category><category>speech-to-text-kostenlos</category><category>ausschussreduzierung</category>
  </item>

    </channel>
  </rss>
