{"id":21287,"date":"2026-10-11T15:33:00","date_gmt":"2026-10-11T13:33:00","guid":{"rendered":"https:\/\/www.lukaswojcik.com\/blog\/?p=21287"},"modified":"2026-10-05T16:49:59","modified_gmt":"2026-10-05T14:49:59","slug":"llm-quantisierung-kontext-modellpassung-de","status":"publish","type":"post","link":"https:\/\/www.lukaswojcik.com\/blog\/de\/cloud-ai-de\/llm-quantisierung-kontext-modellpassung-de\/","title":{"rendered":"Q4, Q8, FP16 und langer Kontext: Warum ein lokales LLM trotz passender Dateigr\u00f6\u00dfe scheitert"},"content":{"rendered":"<p>Eine Modelldatei mit 20 GB passt nicht automatisch auf eine Grafikkarte mit 24 GB. Der Download beschreibt gespeicherte Gewichte und Metadaten. W\u00e4hrend der Ausf\u00fchrung kommen Kontextzustand, tempor\u00e4re Puffer und weitere Belegungen hinzu. Die freie Kapazit\u00e4t ist au\u00dferdem kleiner als der beworbene Gesamtspeicher, sobald andere Anwendungen die GPU verwenden.<\/p>\n<p>Quantisierung und Kontext geh\u00f6ren deshalb in dieselbe Planung. Der Vergleich erkl\u00e4rt Speichermechanismen und Rechenbeispiele; gemessene Qualit\u00e4ts- oder Geschwindigkeitsrangfolgen werden daraus nicht abgeleitet.<\/p>\n<figure class=\"lw-diagram\"><img src=\"https:\/\/www.lukaswojcik.com\/blog\/wp-content\/uploads\/diagrams\/llm-quantisierung-kontext-modellpassung-de.png\" width=\"1120\" height=\"580\" loading=\"lazy\" decoding=\"async\" alt=\"Mehr als die Modelldatei: Gewichte; KV \/ Modellzustand; Laufzeitpuffer; Andere Prozesse; Freie Reserve\"><figcaption>1 GiB Cache im angenommenen Beispiel; keine universelle Modellformel.<\/figcaption><\/figure>\n<h2>Q4, Q8 und FP16 beschreiben verschiedene Budgets<\/h2>\n<table>\n<thead>\n<tr>\n<th>Idealisiertes Beispiel<\/th>\n<th>Rechnung f\u00fcr 30 Milliarden Gewichte<\/th>\n<th>Reiner Gewichtsspeicher<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>4 Bit<\/td>\n<td>30 Milliarden \u00d7 0,5 Byte<\/td>\n<td>15 GB<\/td>\n<\/tr>\n<tr>\n<td>8 Bit<\/td>\n<td>30 Milliarden \u00d7 1 Byte<\/td>\n<td>30 GB<\/td>\n<\/tr>\n<tr>\n<td>16 Bit<\/td>\n<td>30 Milliarden \u00d7 2 Byte<\/td>\n<td>60 GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>Diese Zahlen sind theoretische Dezimalwerte, keine Dateigr\u00f6\u00dfen eines bestimmten Modells. Skalen, Bl\u00f6cke, gemischte Pr\u00e4zisionen und nicht quantisierte Bestandteile ver\u00e4ndern das Ergebnis. GB und GiB sind ebenfalls verschieden: 15 Milliarden Byte entsprechen ungef\u00e4hr 13,97 GiB. Eine Q4-Datei besteht nicht zwingend aus exakt vier Bit f\u00fcr jeden Gesamtparameter.<\/p>\n<h2>Kontext kostet w\u00e4hrend der Ausf\u00fchrung<\/h2>\n<p>Bei klassischer Attention h\u00e4lt der KV-Cache fr\u00fchere Schl\u00fcssel und Werte bereit. F\u00fcr einen vereinfachten dichten Cache lautet die Rechnung: zwei \u00d7 Schichten \u00d7 KV-K\u00f6pfe \u00d7 Kopfbreite \u00d7 Token \u00d7 Byte je Wert \u00d7 Sequenzen. Die Formel gilt nur f\u00fcr die passenden Architektur- und Laufzeitannahmen.<\/p>\n<p>Ein angenommenes Beispiel mit 32 Schichten, acht KV-K\u00f6pfen, Kopfbreite 128, 8.192 Token und zwei Byte je Wert ben\u00f6tigt 1 GiB f\u00fcr eine Sequenz. Vier vollst\u00e4ndig belegte gleichartige Sequenzen ergeben 4 GiB. Gewichte und Laufzeitpuffer kommen zus\u00e4tzlich hinzu. Paging, Cache-Quantisierung oder andere Attention-Muster \u00e4ndern die praktische Belegung.<\/p>\n<h2>Hybride Architekturen brauchen eigene Profile<\/h2>\n<p>Qwen3.5 verwendet eine hybride Struktur mit unterschiedlichen Schichttypen. Eine Formel f\u00fcr ausschlie\u00dflich klassische Attention darf daher nicht ungepr\u00fcft auf jede Schicht angewandt werden. Recurrent-Zustand, Attention-Cache und multimodale Bestandteile ben\u00f6tigen passende Daten aus Modellkonfiguration und Runtime.<\/p>\n<p>Auch MoE trennt aktiven Rechenaufwand vom gesamten Gewichtsspeicher. Eine kleine Zahl aktiver Parameter ist keine Zusage einer kleinen Modelldatei. Speicherplanung anhand des Namens allein scheitert an solchen Unterschieden.<\/p>\n<h2>Auslagerung ver\u00e4ndert das Einsatzprofil<\/h2>\n<p>llama.cpp unterst\u00fctzt hybride CPU-\/GPU-Ausf\u00fchrung. Ein teilweise ausgelagertes Modell kann dadurch \u00fcberhaupt starten. Das ist jedoch eine andere Konfiguration als vollst\u00e4ndige GPU-Ausf\u00fchrung. Arbeitsspeicher, CPU-Leistung und Datentransfers beeinflussen das Antwortverhalten.<\/p>\n<p>Der Vergleich sollte GPU-Belegung, System-RAM, First-Token-Zeit und Ausgabegeschwindigkeit gemeinsam erfassen. Ein erfolgreicher Start ist nur die erste Pr\u00fcfung. Lange Dokumente und mehrere gleichzeitige Sitzungen k\u00f6nnen sp\u00e4ter eine Grenze erreichen, die ein kurzer Chat verdeckt.<\/p>\n<h2>Qualit\u00e4t vor Speicheroptimierung pr\u00fcfen<\/h2>\n<p>Ein gr\u00f6\u00dferes Q4-Modell gewinnt nicht automatisch gegen ein kleineres Q8-Modell. Modelltraining, Architektur und Aufgabe bleiben entscheidend. Derselbe Datensatz mit erwarteten Antworten zeigt Fehler, Auslassungen und Formatprobleme; Dateigr\u00f6\u00dfe allein erkl\u00e4rt die Nutzbarkeit nicht.<\/p>\n<p>Der vorhandene Speicherrechner liefert eine Vorauswahl. Konkrete Modellrevision und Backend, freie Reserve, Kontext und Parallelit\u00e4t geh\u00f6ren in die Eingaben. Die endg\u00fcltige Konfiguration folgt einer realen Speicherpr\u00fcfung und einem kurzen Qualit\u00e4tstest. Eine unbekannte Architektur erh\u00e4lt eine sichtbare Sch\u00e4tzgrenze statt einer scheinbar exakten Zahl.<\/p>\n<p><a href=\"https:\/\/www.lukaswojcik.com\/blog\/de\/werkzeuge\/llm-ollama-vram-memory-estimator\/\">Speicherbedarf f\u00fcr die Modellwahl: LLM-VRAM-Rechner.<\/a><\/p>\n<section class=\"lw-ai-tool-links\">\n<h2>Passende Werkzeuge<\/h2>\n<ul>\n<li><a href=\"https:\/\/www.lukaswojcik.com\/blog\/de\/werkzeuge\/llm-ollama-vram-memory-estimator\/\">LLM-Speicher- und Hardwareplaner<\/a><\/li>\n<\/ul>\n<\/section>\n<div class=\"lw-faq\">\n<h2>Fragen und Antworten<\/h2>\n<h3>Warum passt eine 20-GB-Modelldatei nicht immer auf eine 24-GB-GPU?<\/h3>\n<p>Zur Datei kommen Laufzeitpuffer, Modellzustand beziehungsweise KV-Cache und weiterer Speicherbedarf hinzu. Die tats\u00e4chliche Belegung h\u00e4ngt von Runtime, Architektur, Kontext und Parallelit\u00e4t ab.<\/p>\n<h3>Ist ein gr\u00f6\u00dferes Q4-Modell immer besser als ein kleineres Q8-Modell?<\/h3>\n<p>Die Pr\u00e4zision allein liefert keine Qualit\u00e4tsrangfolge. Modellfamilie, Training und Aufgabe bleiben entscheidend. Ein Vergleich ben\u00f6tigt gleiche Aufgaben, dokumentierte Konfigurationen und eine Pr\u00fcfung der Fehler statt nur der Dateigr\u00f6\u00dfe.<\/p>\n<\/div>\n<h2>Quellen<\/h2>\n<ol>\n<li id=\"source-1\"><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\" target=\"_blank\" rel=\"noopener noreferrer\">llama.cpp: Projekt und Backends<\/a><\/li>\n<li id=\"source-2\"><a href=\"https:\/\/huggingface.co\/Qwen\/Qwen3.5-9B\" target=\"_blank\" rel=\"noopener noreferrer\">Qwen: Qwen3.5-9B, offizielle Model Card<\/a><\/li>\n<li id=\"source-3\"><a href=\"https:\/\/huggingface.co\/mistralai\/Ministral-3-14B-Instruct-2512\" target=\"_blank\" rel=\"noopener noreferrer\">Mistral: Ministral 3 14B Instruct, offizielle Model Card<\/a><\/li>\n<li id=\"source-4\"><a href=\"https:\/\/docs.vllm.ai\/en\/latest\/getting_started\/installation\/gpu\/\" target=\"_blank\" rel=\"noopener noreferrer\">vLLM: GPU installation<\/a><\/li>\n<\/ol>\n<p><small>Quellen gepr\u00fcft: 5. Oktober 2026.<\/small><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Gewichte, Cache und Laufzeitreserve gemeinsam planen: Q4, Q8, FP16, lange Kontexte und hybride Architekturen im Vergleich.<\/p>\n","protected":false},"author":1,"featured_media":21395,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[104],"tags":[],"class_list":["post-21287","post","type-post","status-publish","format-standard","hentry","category-cloud-ai-de"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.1 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Q4, Q8, FP16 und langer Kontext: Warum ein lokales LLM trotz passender Dateigr\u00f6\u00dfe scheitert | Lukas Wojcik<\/title>\n<meta name=\"description\" content=\"Gewichte, Cache und Laufzeitreserve gemeinsam planen: Q4, Q8, FP16, lange Kontexte und hybride Architekturen im Vergleich.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/www.lukaswojcik.com\/blog\/de\/cloud-ai-de\/llm-quantisierung-kontext-modellpassung-de\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Q4, Q8, FP16 und langer Kontext: Warum ein lokales LLM trotz passender Dateigr\u00f6\u00dfe scheitert | Lukas Wojcik\" \/>\n<meta property=\"og:description\" content=\"Gewichte, Cache und Laufzeitreserve gemeinsam planen: Q4, Q8, FP16, lange Kontexte und hybride Architekturen im Vergleich.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/www.lukaswojcik.com\/blog\/de\/cloud-ai-de\/llm-quantisierung-kontext-modellpassung-de\/\" \/>\n<meta property=\"og:site_name\" content=\"Lukas Wojcik - Blog\" \/>\n<meta property=\"article:published_time\" content=\"2026-10-11T13:33:00+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/www.lukaswojcik.com\/blog\/wp-content\/uploads\/2026\/10\/hero-21287-cloudai20261005-dark.png\" \/>\n\t<meta property=\"og:image:width\" content=\"1200\" \/>\n\t<meta property=\"og:image:height\" content=\"630\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"Lukas Wojcik\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Lukas Wojcik\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"3 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/de\\\/cloud-ai-de\\\/llm-quantisierung-kontext-modellpassung-de\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/de\\\/cloud-ai-de\\\/llm-quantisierung-kontext-modellpassung-de\\\/\"},\"author\":{\"name\":\"Lukas Wojcik\",\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/#\\\/schema\\\/person\\\/895f7604f9b6b71aad9bba33af28d0f9\"},\"headline\":\"Q4, Q8, FP16 und langer Kontext: Warum ein lokales LLM trotz passender Dateigr\u00f6\u00dfe scheitert\",\"datePublished\":\"2026-10-11T13:33:00+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/de\\\/cloud-ai-de\\\/llm-quantisierung-kontext-modellpassung-de\\\/\"},\"wordCount\":636,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/#\\\/schema\\\/person\\\/895f7604f9b6b71aad9bba33af28d0f9\"},\"image\":{\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/de\\\/cloud-ai-de\\\/llm-quantisierung-kontext-modellpassung-de\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/10\\\/hero-21287-cloudai20261005-dark.png\",\"articleSection\":[\"Cloud &amp; AI\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/de\\\/cloud-ai-de\\\/llm-quantisierung-kontext-modellpassung-de\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/de\\\/cloud-ai-de\\\/llm-quantisierung-kontext-modellpassung-de\\\/\",\"url\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/de\\\/cloud-ai-de\\\/llm-quantisierung-kontext-modellpassung-de\\\/\",\"name\":\"Q4, Q8, FP16 und langer Kontext: Warum ein lokales LLM trotz passender Dateigr\u00f6\u00dfe scheitert | Lukas Wojcik\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/de\\\/cloud-ai-de\\\/llm-quantisierung-kontext-modellpassung-de\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/de\\\/cloud-ai-de\\\/llm-quantisierung-kontext-modellpassung-de\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/10\\\/hero-21287-cloudai20261005-dark.png\",\"datePublished\":\"2026-10-11T13:33:00+00:00\",\"description\":\"Gewichte, Cache und Laufzeitreserve gemeinsam planen: Q4, Q8, FP16, lange Kontexte und hybride Architekturen im Vergleich.\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/de\\\/cloud-ai-de\\\/llm-quantisierung-kontext-modellpassung-de\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/de\\\/cloud-ai-de\\\/llm-quantisierung-kontext-modellpassung-de\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/de\\\/cloud-ai-de\\\/llm-quantisierung-kontext-modellpassung-de\\\/#primaryimage\",\"url\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/10\\\/hero-21287-cloudai20261005-dark.png\",\"contentUrl\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/10\\\/hero-21287-cloudai20261005-dark.png\",\"width\":1200,\"height\":630,\"caption\":\"Q4, Q8, FP16 und langer Kontext: Warum ein lokales LLM trotz passender Dateigr\u00f6\u00dfe scheitert\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/de\\\/cloud-ai-de\\\/llm-quantisierung-kontext-modellpassung-de\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Q4, Q8, FP16 und langer Kontext: Warum ein lokales LLM trotz passender Dateigr\u00f6\u00dfe scheitert\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/#website\",\"url\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/\",\"name\":\"Lukas Wojcik - Blog\",\"description\":\"\",\"publisher\":{\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/#\\\/schema\\\/person\\\/895f7604f9b6b71aad9bba33af28d0f9\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":[\"Person\",\"Organization\"],\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/#\\\/schema\\\/person\\\/895f7604f9b6b71aad9bba33af28d0f9\",\"name\":\"Lukas Wojcik\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/lw-x2.jpg\",\"url\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/lw-x2.jpg\",\"contentUrl\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/lw-x2.jpg\",\"width\":424,\"height\":636,\"caption\":\"Lukas Wojcik\"},\"logo\":{\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/lw-x2.jpg\"},\"sameAs\":[\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\"]}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Q4, Q8, FP16 und langer Kontext: Warum ein lokales LLM trotz passender Dateigr\u00f6\u00dfe scheitert | Lukas Wojcik","description":"Gewichte, Cache und Laufzeitreserve gemeinsam planen: Q4, Q8, FP16, lange Kontexte und hybride Architekturen im Vergleich.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/www.lukaswojcik.com\/blog\/de\/cloud-ai-de\/llm-quantisierung-kontext-modellpassung-de\/","og_locale":"en_US","og_type":"article","og_title":"Q4, Q8, FP16 und langer Kontext: Warum ein lokales LLM trotz passender Dateigr\u00f6\u00dfe scheitert | Lukas Wojcik","og_description":"Gewichte, Cache und Laufzeitreserve gemeinsam planen: Q4, Q8, FP16, lange Kontexte und hybride Architekturen im Vergleich.","og_url":"https:\/\/www.lukaswojcik.com\/blog\/de\/cloud-ai-de\/llm-quantisierung-kontext-modellpassung-de\/","og_site_name":"Lukas Wojcik - Blog","article_published_time":"2026-10-11T13:33:00+00:00","og_image":[{"width":1200,"height":630,"url":"https:\/\/www.lukaswojcik.com\/blog\/wp-content\/uploads\/2026\/10\/hero-21287-cloudai20261005-dark.png","type":"image\/png"}],"author":"Lukas Wojcik","twitter_card":"summary_large_image","twitter_misc":{"Written by":"Lukas Wojcik","Est. reading time":"3 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/www.lukaswojcik.com\/blog\/de\/cloud-ai-de\/llm-quantisierung-kontext-modellpassung-de\/#article","isPartOf":{"@id":"https:\/\/www.lukaswojcik.com\/blog\/de\/cloud-ai-de\/llm-quantisierung-kontext-modellpassung-de\/"},"author":{"name":"Lukas Wojcik","@id":"https:\/\/www.lukaswojcik.com\/blog\/#\/schema\/person\/895f7604f9b6b71aad9bba33af28d0f9"},"headline":"Q4, Q8, FP16 und langer Kontext: Warum ein lokales LLM trotz passender Dateigr\u00f6\u00dfe scheitert","datePublished":"2026-10-11T13:33:00+00:00","mainEntityOfPage":{"@id":"https:\/\/www.lukaswojcik.com\/blog\/de\/cloud-ai-de\/llm-quantisierung-kontext-modellpassung-de\/"},"wordCount":636,"commentCount":0,"publisher":{"@id":"https:\/\/www.lukaswojcik.com\/blog\/#\/schema\/person\/895f7604f9b6b71aad9bba33af28d0f9"},"image":{"@id":"https:\/\/www.lukaswojcik.com\/blog\/de\/cloud-ai-de\/llm-quantisierung-kontext-modellpassung-de\/#primaryimage"},"thumbnailUrl":"https:\/\/www.lukaswojcik.com\/blog\/wp-content\/uploads\/2026\/10\/hero-21287-cloudai20261005-dark.png","articleSection":["Cloud &amp; AI"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/www.lukaswojcik.com\/blog\/de\/cloud-ai-de\/llm-quantisierung-kontext-modellpassung-de\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/www.lukaswojcik.com\/blog\/de\/cloud-ai-de\/llm-quantisierung-kontext-modellpassung-de\/","url":"https:\/\/www.lukaswojcik.com\/blog\/de\/cloud-ai-de\/llm-quantisierung-kontext-modellpassung-de\/","name":"Q4, Q8, FP16 und langer Kontext: Warum ein lokales LLM trotz passender Dateigr\u00f6\u00dfe scheitert | Lukas Wojcik","isPartOf":{"@id":"https:\/\/www.lukaswojcik.com\/blog\/#website"},"primaryImageOfPage":{"@id":"https:\/\/www.lukaswojcik.com\/blog\/de\/cloud-ai-de\/llm-quantisierung-kontext-modellpassung-de\/#primaryimage"},"image":{"@id":"https:\/\/www.lukaswojcik.com\/blog\/de\/cloud-ai-de\/llm-quantisierung-kontext-modellpassung-de\/#primaryimage"},"thumbnailUrl":"https:\/\/www.lukaswojcik.com\/blog\/wp-content\/uploads\/2026\/10\/hero-21287-cloudai20261005-dark.png","datePublished":"2026-10-11T13:33:00+00:00","description":"Gewichte, Cache und Laufzeitreserve gemeinsam planen: Q4, Q8, FP16, lange Kontexte und hybride Architekturen im Vergleich.","breadcrumb":{"@id":"https:\/\/www.lukaswojcik.com\/blog\/de\/cloud-ai-de\/llm-quantisierung-kontext-modellpassung-de\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/www.lukaswojcik.com\/blog\/de\/cloud-ai-de\/llm-quantisierung-kontext-modellpassung-de\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.lukaswojcik.com\/blog\/de\/cloud-ai-de\/llm-quantisierung-kontext-modellpassung-de\/#primaryimage","url":"https:\/\/www.lukaswojcik.com\/blog\/wp-content\/uploads\/2026\/10\/hero-21287-cloudai20261005-dark.png","contentUrl":"https:\/\/www.lukaswojcik.com\/blog\/wp-content\/uploads\/2026\/10\/hero-21287-cloudai20261005-dark.png","width":1200,"height":630,"caption":"Q4, Q8, FP16 und langer Kontext: Warum ein lokales LLM trotz passender Dateigr\u00f6\u00dfe scheitert"},{"@type":"BreadcrumbList","@id":"https:\/\/www.lukaswojcik.com\/blog\/de\/cloud-ai-de\/llm-quantisierung-kontext-modellpassung-de\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/www.lukaswojcik.com\/blog\/"},{"@type":"ListItem","position":2,"name":"Q4, Q8, FP16 und langer Kontext: Warum ein lokales LLM trotz passender Dateigr\u00f6\u00dfe scheitert"}]},{"@type":"WebSite","@id":"https:\/\/www.lukaswojcik.com\/blog\/#website","url":"https:\/\/www.lukaswojcik.com\/blog\/","name":"Lukas Wojcik - Blog","description":"","publisher":{"@id":"https:\/\/www.lukaswojcik.com\/blog\/#\/schema\/person\/895f7604f9b6b71aad9bba33af28d0f9"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/www.lukaswojcik.com\/blog\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":["Person","Organization"],"@id":"https:\/\/www.lukaswojcik.com\/blog\/#\/schema\/person\/895f7604f9b6b71aad9bba33af28d0f9","name":"Lukas Wojcik","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.lukaswojcik.com\/blog\/wp-content\/uploads\/2026\/07\/lw-x2.jpg","url":"https:\/\/www.lukaswojcik.com\/blog\/wp-content\/uploads\/2026\/07\/lw-x2.jpg","contentUrl":"https:\/\/www.lukaswojcik.com\/blog\/wp-content\/uploads\/2026\/07\/lw-x2.jpg","width":424,"height":636,"caption":"Lukas Wojcik"},"logo":{"@id":"https:\/\/www.lukaswojcik.com\/blog\/wp-content\/uploads\/2026\/07\/lw-x2.jpg"},"sameAs":["https:\/\/www.lukaswojcik.com\/blog"]}]}},"_links":{"self":[{"href":"https:\/\/www.lukaswojcik.com\/blog\/wp-json\/wp\/v2\/posts\/21287","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.lukaswojcik.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.lukaswojcik.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.lukaswojcik.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.lukaswojcik.com\/blog\/wp-json\/wp\/v2\/comments?post=21287"}],"version-history":[{"count":1,"href":"https:\/\/www.lukaswojcik.com\/blog\/wp-json\/wp\/v2\/posts\/21287\/revisions"}],"predecessor-version":[{"id":21359,"href":"https:\/\/www.lukaswojcik.com\/blog\/wp-json\/wp\/v2\/posts\/21287\/revisions\/21359"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.lukaswojcik.com\/blog\/wp-json\/wp\/v2\/media\/21395"}],"wp:attachment":[{"href":"https:\/\/www.lukaswojcik.com\/blog\/wp-json\/wp\/v2\/media?parent=21287"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.lukaswojcik.com\/blog\/wp-json\/wp\/v2\/categories?post=21287"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.lukaswojcik.com\/blog\/wp-json\/wp\/v2\/tags?post=21287"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}