# GLM-5.2 Fast Preview > Low-latency chat completions from the GLM-5.2 weights on throughput-tuned serving — consistently faster than the standard tier, at twice the price and identical answer quality. ## Overview - **Endpoint**: `https://queue.modelrunner.run/z-ai/glm-5.2-fast-preview` - **Model ID**: `z-ai/glm-5.2-fast-preview` - **Category**: text-to-text - **Kind**: inference - **Tags**: glm, glm-5.2, glm-5.2-fast-preview, fast, z-ai, zhipu, llm, text-to-text, chat, chat-completions, openai-compatible, low-latency, fast-inference, real-time, high-throughput, streaming, reasoning, thinking, agentic, agentic-coding, coding, code-generation, long-context, open-weight, open-source, mit-license, tool-calling, json-mode, preview ## Pricing - **Input tokens**: $2.8 per 1M - **Cached input tokens**: $0.7 per 1M - **Output tokens**: $8.8 per 1M ## Request Lifecycle This model runs on the ModelRunner **asynchronous queue API** — a single POST does not return the output. Every call requires an `Authorization: Key $MODEL_RUNNER_KEY` header. Run three steps: 1. **Submit** — `POST https://queue.modelrunner.run/z-ai/glm-5.2-fast-preview` with a JSON body holding the input fields at the top level. The body may also include a reserved top-level `metadata` object — a flat string map (max 16 keys, key ≤64 / value ≤512 chars) stored on the request for your own tagging. It is never sent to the model; filter your request history with `GET https://queue.modelrunner.run/requests?metadata=` (exact key=value matches, AND-ed). The response carries request handles only (no output yet): ```json { "status": "IN_QUEUE", "request_id": "<21-char id>", "status_url": "https://queue.modelrunner.run/z-ai/glm-5.2-fast-preview/requests//status", "response_url": "https://queue.modelrunner.run/z-ai/glm-5.2-fast-preview/requests/", "cancel_url": "https://queue.modelrunner.run/z-ai/glm-5.2-fast-preview/requests//cancel" } ``` 2. **Poll status** — `GET ` until `status` is `COMPLETED`. Possible values are `IN_QUEUE`, `IN_PROGRESS`, `COMPLETED`, `FAILED`, `CANCELLED`. A `FAILED` request responds with HTTP 400 and an `error` field. 3. **Read result** — `GET `. Returns the finished request, including the generated `output`: ```json { "id": "", "status": "COMPLETED", "output": ..., "input": ... } ``` The JavaScript and Python SDKs below perform steps 2–3 for you. In any language without an SDK (Swift, Go, Kotlin, etc.) you must implement the polling loop and the final result fetch yourself — see the cURL example for the full flow. ### Input Schema - **`seed`** (`integer`, _optional_): Best-effort determinism hint. - **`stop`** (`unknown`, _optional_): Up to 4 stop sequences. - **`tools`** (`array`, _optional_): OpenAI-format tool definitions the model may call. - **`stream`** (`boolean`, _optional_): Return the reply as a Server-Sent Events stream of deltas terminated by \`data: \[DONE\]\`. - Default: `false` - **`messages`** (`array`, _required_): OpenAI-style conversation history. Each item is an object with a \`role\` (\`system\`, \`user\`, \`assistant\` or \`tool\`) and \`content\`. - **`max_tokens`** (`integer`, _optional_): Upper bound on generated tokens, up to the family's published 131,072-token output ceiling. Thinking consumes this budget, so allow generous headroom. - Range: `1` to `"+inf"` - **`tool_choice`** (`unknown`, _optional_): \`auto\`, \`none\`, \`required\`, or a specific tool. - **`response_format`** (`object`, _optional_): Structured-output control. Set its \`type\` to \`json_object\` to force a JSON reply. Structured output is supported in non-thinking mode only, so pair it with \`reasoning_effort\` set to \`none\`. - **`reasoning_effort`** (`ReasoningEffortEnum`, _optional_): How hard the model thinks before answering, across seven levels. Defaults to \`max\`, the highest — lower it to cut cost, because thinking tokens bill as output tokens. \`none\` disables reasoning entirely. - Default: `"max"` - Options: `"none"`, `"minimal"`, `"low"`, `"medium"`, `"high"`, `"xhigh"`, `"max"` ### Output Schema - **`id`** (`string`, _optional_): - **`model`** (`string`, _optional_): - **`usage`** (`object`, _optional_): - **`object`** (`string`, _optional_): - **`choices`** (`array`, _optional_): - **`created`** (`integer`, _optional_): ## Default Example **Input** ```json { "messages": [ { "role": "user", "content": "Here is a Python function that should return the k most frequent words in a text, with ties broken alphabetically:\n\ndef top_k_words(text, k):\n words = text.lower().split()\n counts = {}\n for w in words:\n counts[w] = counts.get(w, 0) + 1\n ranked = sorted(counts.items(), key=lambda kv: -kv[1])\n return [w for w, _ in ranked[:k]]\n\nFind every bug, explain why each one is wrong, then give a corrected version with its time and space complexity." } ], "max_tokens": 8000, "reasoning_effort": "high" } ``` **Output** ```json { "id": "chatcmpl-Xh0VPXLvzJFy2E4NR6h4z", "model": "z-ai/glm-5.2-fast-preview", "usage": { "total_tokens": 2102, "prompt_tokens": 130, "completion_tokens": 1972, "prompt_tokens_details": { "cached_tokens": 0 }, "completion_tokens_details": { "reasoning_tokens": 1261 } }, "object": "chat.completion", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Here are the bugs in the provided function:\n\n### Bug 1: Ties are not broken alphabetically\n**Why it's wrong:** The sorting key `lambda kv: -kv[1]` only sorts the words by their frequency in descending order. Python's `sorted` function is stable, meaning that if two words have the same frequency, they will retain their original relative order (which is the order they were inserted into the dictionary), not alphabetical order. \n**Fix:** Change the sorting key to a tuple: `lambda kv: (-kv[1], kv[0])`. This sorts by frequency descending (`-kv[1]`), and for ties, it sorts by the word alphabetically ascending (`kv[0]`).\n\n### Bug 2: Punctuation is not stripped\n**Why it's wrong:** The function uses `text.lower().split()`, which splits the text by whitespace but leaves punctuation attached to the words. For example, the word `\"apple\"` and `\"apple!\"` will be treated as two completely different words, resulting in incorrect frequency counts.\n**Fix:** Use a regular expression like `re.findall(r'\\w+', text.lower())` to extract only the alphanumeric words, ignoring punctuation.\n\n### Bug 3 (Edge Case): No handling for negative `k`\n**Why it's wrong:** If `k` is negative, `ranked[:k]` behaves unexpectedly in Python (it returns all items except the last `k` items). A top-k function should ideally just return an empty list if `k <= 0`.\n**Fix:** Add an early return `if k <= 0: return []`.\n\n### Corrected Version\n\n```python\nimport re\n\ndef top_k_words(text, k):\n if k <= 0:\n return []\n \n # Extract words (alphanumeric sequences) and convert to lowercase\n words = re.findall(r'\\w+', text.lower())\n \n # Count frequencies\n counts = {}\n for w in words:\n counts[w] = counts.get(w, 0) + 1\n \n # Sort by frequency (descending) and alphabetically (ascending) for ties\n ranked = sorted(counts.items(), key=lambda kv: (-kv[1], kv[0]))\n \n return [w for w, _ in ranked[:k]]\n```\n\n### Complexity\nLet \\(N\\) be the number of words in the text, and \\(U\\) be the number of unique words (so \\(U \\le N\\)).\n\n**Time Complexity:** \n- Finding words and counting: \\(O(N)\\) \n- Sorting the unique words: \\(O(U \\log U)\\)\n- Slicing the list: \\(O(k)\\)\nOverall Time Complexity: **\\(O(N + U \\log U)\\)**\n\n*(Note: If we wanted to optimize the time complexity strictly for \\(k\\), we could use a min-heap of size \\(k\\) via `heapq.nlargest`, which would reduce the time complexity to \\(O(N + U \\log k)\\). However, the standard `sorted` approach is cleaner and usually sufficient unless \\(U\\) is massive).*\n\n**Space Complexity:** \n- The `counts` dictionary takes \\(O(U)\\) space.\n- The `ranked` list takes \\(O(U)\\) space.\nOverall Space Complexity: **\\(O(U)\\)** (or \\(O(N)\\) in the worst case where every word is unique).", "reasoning_content": "The user wants to find bugs in the provided Python function, explain them, and provide a corrected version along with time and space complexity.\n\nFunction:\n```python\ndef top_k_words(text, k):\n words = text.lower().split()\n counts = {}\n for w in words:\n counts[w] = counts.get(w, 0) + 1\n ranked = sorted(counts.items(), key=lambda kv: -kv[1])\n return [w for w, _ in ranked[:k]]\n```\n\nThe problem statement says: \"return the k most frequent words in a text, with ties broken alphabetically\".\nCurrently, the sorting key is `lambda kv: -kv[1]`. This sorts by frequency descending. If there are ties in frequency, the order depends on the original dictionary order (insertion order in Python 3.7+), not alphabetically.\nTo break ties alphabetically, we need the words themselves to be sorted in ascending alphabetical order.\nSo, the sorting key should be `lambda kv: (-kv[1], kv[0])`.\n\nLet's check for other bugs.\n`words = text.lower().split()`\nWait, `split()` splits by whitespace, but what about punctuation? \"Hello, world!\" -> \"hello,\" and \"world!\". Often, word counting requires stripping punctuation. Does the prompt explicitly say we need to handle punctuation? \"k most frequent words in a text\". Usually, basic string processing requires at least handling punctuation, but maybe not strictly required unless specified. Let's see if there's a simpler bug.\nTies broken alphabetically: `-kv[1]` only handles frequency. So the missing alphabetical tie-breaking is the primary bug.\nWait, another bug: if `k` is greater than the number of unique words, `ranked[:k]` works fine.\nAre there any other bugs?\nUsing `-kv[1]` is fine but using `(-kv[1], kv[0])` is better.\nWhat about stripping punctuation? \"text.lower().split()\" might include punctuation attached to words. A more robust approach uses `re.findall(r'\\b\\w+\\b', text.lower())` or `re.sub(r'[^\\w\\s]', '', text.lower()).split()`. I should mention punctuation as a potential issue or bug depending on the strict definition of \"word\". But let's focus on the tie-breaking first.\n\nBug 1: Ties are not broken alphabetically. The current sorting key `lambda kv: -kv[1]` only sorts by frequency in descending order. Python's `sorted` is stable, meaning that if two items have the same sorting key (which they will if they have the same frequency), they retain their original relative order, which is the insertion order into the dictionary, not alphabetical order. To fix this, the sorting key should be `lambda kv: (-kv[1], kv[0])`. This sorts by frequency descending, and then by word alphabetically ascending.\n\nBug 2 (Potential): Punctuation is not handled. Words like \"word\" and \"word.\" will be treated as two different words because `split()` only removes whitespace. To fix this, we should strip punctuation or use a regular expression to extract words. E.g., `re.findall(r'\\b[a-zA-Z]+\\b', text.lower())` or `re.findall(r'\\w+', text.lower())`.\n\nLet's write down the bugs clearly.\n\n1. Tie-breaking logic is missing.\n2. Punctuation is not stripped.\n\nWait, the prompt says \"Find every bug\".\nIs it possible that the `-` in `-kv[1]` is considered bad practice or bug? It's not a bug, but `reverse=True` is preferred. Not a bug.\n\nLet's construct the response.\n\nBug 1: Missing alphabetical tie-breaking.\nBug 2: Punctuation is not removed. (Or at least it's an edge case that typically constitutes a bug in word count problems).\n\nLet's provide the corrected version.\n\n```python\nimport re\n\ndef top_k_words(text, k):\n # Convert to lowercase and find all words (alphanumeric sequences)\n words = re.findall(r'\\w+', text.lower())\n counts = {}\n for w in words:\n counts[w] = counts.get(w, 0) + 1\n # Sort by frequency descending, then alphabetically ascending\n ranked = sorted(counts.items(), key=lambda kv: (-kv[1], kv[0]))\n return [w for w, _ in ranked[:k]]\n```\n\nTime complexity: O(N log N) where N is the number of unique words. If we use a heap, we can do O(N log k). The current `sorted` takes O(N log N).\nLet's mention both. O(N log N) for the `sorted` approach. If we use `heapq.nlargest`, it's O(N log k).\nSpace complexity: O(N) to store the dictionary and the list of unique words.\n\nLet's refine the explanation.\nBug 1: `key=lambda kv: -kv[1]` doesn't handle ties alphabetically. `sorted` is stable, so ties remain in insertion order. Fix: `key=lambda kv: (-kv[1], kv[0])`.\nBug 2: `split()` doesn't remove punctuation. \"word.\" and \"word\" will be counted as two separate words. Fix: use `re.findall(r'\\w+', text)` or similar.\n\nAre there any other bugs?\nWhat if `k` is 0 or negative? `ranked[:k]` handles `k=0` (returns empty list). If `k < 0`, `ranked[:-1]` returns all but the last, which is wrong. If the prompt says \"k most frequent\", usually k is positive, but handling `k <= 0` by returning `[]` is safer. I will mention it but it's minor. Let's assume k >= 0.\n\nLet's structure the answer:\n- Bug 1: Tie-breaking\n- Bug 2: Punctuation\n- Corrected Version\n- Complexity" }, "finish_reason": "stop" } ], "created": 1787136499 } ``` ## Usage Examples ### cURL The queue API is asynchronous: submit the request, poll `status_url` until it is `COMPLETED`, then read the result from `response_url`. Requires `jq`. ```bash # 1. Submit the request (returns request handles, not the output) SUBMIT=$(curl --silent --request POST \ --url https://queue.modelrunner.run/z-ai/glm-5.2-fast-preview \ --header "Authorization: Key $MODEL_RUNNER_KEY" \ --header "Content-Type: application/json" \ --data '{ "messages": [ { "role": "user", "content": "Here is a Python function that should return the k most frequent words in a text, with ties broken alphabetically:\n\ndef top_k_words(text, k):\n words = text.lower().split()\n counts = {}\n for w in words:\n counts[w] = counts.get(w, 0) + 1\n ranked = sorted(counts.items(), key=lambda kv: -kv[1])\n return [w for w, _ in ranked[:k]]\n\nFind every bug, explain why each one is wrong, then give a corrected version with its time and space complexity." } ], "max_tokens": 8000, "reasoning_effort": "high" }') STATUS_URL=$(echo "$SUBMIT" | jq -r '.status_url') RESPONSE_URL=$(echo "$SUBMIT" | jq -r '.response_url') # 2. Poll until the request leaves the queue / in-progress state while true; do STATUS=$(curl --silent --url "$STATUS_URL" \ --header "Authorization: Key $MODEL_RUNNER_KEY" | jq -r '.status') echo "Status: $STATUS" case "$STATUS" in COMPLETED) break ;; FAILED|CANCELLED) echo "Request $STATUS"; exit 1 ;; esac sleep 1 done # 3. Read the finished request, including the generated output curl --silent --url "$RESPONSE_URL" \ --header "Authorization: Key $MODEL_RUNNER_KEY" ``` ### JavaScript ```javascript import { modelrunner } from "@modelrunner/client"; const result = await modelrunner.subscribe("z-ai/glm-5.2-fast-preview", { input: { "messages": [ { "role": "user", "content": "Here is a Python function that should return the k most frequent words in a text, with ties broken alphabetically:\n\ndef top_k_words(text, k):\n words = text.lower().split()\n counts = {}\n for w in words:\n counts[w] = counts.get(w, 0) + 1\n ranked = sorted(counts.items(), key=lambda kv: -kv[1])\n return [w for w, _ in ranked[:k]]\n\nFind every bug, explain why each one is wrong, then give a corrected version with its time and space complexity." } ], "max_tokens": 8000, "reasoning_effort": "high" } }); console.log(result.data); ``` ### Python ```python import asyncio import modelrunner_ai async def main(): response = await modelrunner_ai.submit_async( "z-ai/glm-5.2-fast-preview", arguments={ "messages": [ { "role": "user", "content": "Here is a Python function that should return the k most frequent words in a text, with ties broken alphabetically:\n\ndef top_k_words(text, k):\n words = text.lower().split()\n counts = {}\n for w in words:\n counts[w] = counts.get(w, 0) + 1\n ranked = sorted(counts.items(), key=lambda kv: -kv[1])\n return [w for w, _ in ranked[:k]]\n\nFind every bug, explain why each one is wrong, then give a corrected version with its time and space complexity." } ], "max_tokens": 8000, "reasoning_effort": "high" } ) result = await response.get() print(result["output"]) asyncio.run(main()) ``` ## Additional Resources - [Playground](https://modelrunner.ai/models/z-ai/glm-5.2-fast-preview) - [OpenAPI Schema](https://modelrunner.ai/models/z-ai/glm-5.2-fast-preview/openapi.json) - [LLM Instructions](https://modelrunner.ai/models/z-ai/glm-5.2-fast-preview/llms.txt) - [GitHub](https://github.com/zai-org/GLM-5) - [License](https://huggingface.co/zai-org/GLM-5.2/blob/main/README.md) - [Weights](https://huggingface.co/zai-org/GLM-5.2) - [Paper](https://arxiv.org/abs/2602.15763)