<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>AI |</title><link>https://thecloudadmin.eu/category/ai/</link><atom:link href="https://thecloudadmin.eu/category/ai/index.xml" rel="self" type="application/rss+xml"/><description>AI</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Sun, 11 Oct 2026 09:59:01 +0000</lastBuildDate><image><url>https://thecloudadmin.eu/media/icon_hu_56fcbaee5450a8b1.png</url><title>AI</title><link>https://thecloudadmin.eu/category/ai/</link></image><item><title>Local AI on my Surface Book 3 with Ollama and Qwen3:8b</title><link>https://thecloudadmin.eu/blog/2026/10/local-ai-surface-book-3-ollama-qwen3/</link><pubDate>Sun, 11 Oct 2026 09:59:01 +0000</pubDate><guid>https://thecloudadmin.eu/blog/2026/10/local-ai-surface-book-3-ollama-qwen3/</guid><description>&lt;h2 id="introduction"&gt;Introduction&lt;/h2&gt;
&lt;p&gt;Running AI locally does not have to mean running everything on the laptop you use for daily work. In my setup, I split those roles between two Surface devices.&lt;/p&gt;
&lt;p&gt;My &lt;strong&gt;Surface Book 3&lt;/strong&gt; runs &lt;strong&gt;Ollama&lt;/strong&gt; with &lt;strong&gt;Qwen3:8b&lt;/strong&gt;. My &lt;strong&gt;Surface Laptop 7 for Business&lt;/strong&gt; is where I interact with that model, using both &lt;strong&gt;VS Code Chat&lt;/strong&gt; and &lt;strong&gt;GitHub Copilot CLI&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The interesting part is that the AI workload runs on a different device from my editor and terminal. The Surface Book 3 hosts the model, while the Surface Laptop connects to it over the network.&lt;/p&gt;
&lt;p&gt;This post explains that architecture, the configuration needed to connect the two devices, and what my Ollama logs reveal about performance. The figures come from a small number of real requests, not a controlled benchmark suite.&lt;/p&gt;
&lt;h2 id="one-surface-hosts-the-model-the-other-is-the-client"&gt;One Surface Hosts the Model, the Other Is the Client&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Surface Book 3&lt;/td&gt;
&lt;td&gt;Hosts Ollama and runs model inference&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ollama&lt;/td&gt;
&lt;td&gt;Loads the model and exposes an API&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3:8b&lt;/td&gt;
&lt;td&gt;The local language model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Surface Laptop 7 for Business&lt;/td&gt;
&lt;td&gt;Runs the editor and terminal clients&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;VS Code Chat&lt;/td&gt;
&lt;td&gt;Provides an editor-based chat interface&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GitHub Copilot CLI&lt;/td&gt;
&lt;td&gt;Provides a terminal-based coding-agent interface&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The request path is straightforward:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Surface Laptop 7 for Business
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; VS Code Chat / GitHub Copilot CLI
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; |
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; | Private network connection
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; v
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;Surface Book 3
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; Ollama API -&amp;gt; Qwen3:8b
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Here, &lt;strong&gt;local AI&lt;/strong&gt; means inference on my own hardware. It does not mean that the model runs on the Surface Laptop itself, nor that the entire development workflow is automatically offline.&lt;/p&gt;
&lt;h2 id="preparing-qwen38b-on-the-surface-book-3"&gt;Preparing Qwen3:8b on the Surface Book 3&lt;/h2&gt;
&lt;p&gt;Install
on the Surface Book 3. With Ollama running, download the model:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-shell" data-lang="shell"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ollama pull qwen3:8b
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Then test it directly on the host:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-shell" data-lang="shell"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ollama run qwen3:8b
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Start with a small prompt, such as asking the model to explain a short function. This separates model-loading problems from network or client-configuration problems.&lt;/p&gt;
&lt;p&gt;Use the following commands to inspect the installation and the running model:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-shell" data-lang="shell"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ollama list
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;ollama ps
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;code&gt;ollama list&lt;/code&gt; shows the downloaded models. While the model is loaded, &lt;code&gt;ollama ps&lt;/code&gt; shows whether it is using CPU, GPU, or a combination of both.&lt;/p&gt;
&lt;p&gt;That distinction matters on a Surface Book 3. Hardware configurations differ, and having a GPU does not guarantee that the entire model and its context fit in GPU memory. Check the actual allocation instead of assuming full GPU acceleration.&lt;/p&gt;
&lt;h3 id="my-surface-book-3-configuration"&gt;My Surface Book 3 Configuration&lt;/h3&gt;
&lt;p&gt;The logs from October 6 and 7, 2026 show the following hardware and runtime configuration:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Observed configuration&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;CPU&lt;/td&gt;
&lt;td&gt;Intel Core i7-1065G7 at a reported base frequency of 1.30 GHz&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;System memory&lt;/td&gt;
&lt;td&gt;31.6 GiB reported by Ollama, consistent with a 32 GB configuration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPU&lt;/td&gt;
&lt;td&gt;NVIDIA GeForce GTX 1660 Ti with Max-Q Design&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Integrated graphics&lt;/td&gt;
&lt;td&gt;Intel Iris Plus&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPU memory&lt;/td&gt;
&lt;td&gt;6,196 MiB reported by the backend&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Storage&lt;/td&gt;
&lt;td&gt;Approximately 690 GB free when I set it up&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Operating system&lt;/td&gt;
&lt;td&gt;Windows 11 Enterprise Insider Preview, build 26340&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ollama server version&lt;/td&gt;
&lt;td&gt;0.35.1 in the network setup log&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inference backend&lt;/td&gt;
&lt;td&gt;Vulkan, with CPU participation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model&lt;/td&gt;
&lt;td&gt;Qwen3 8B, 8.19 billion parameters&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quantization&lt;/td&gt;
&lt;td&gt;Q4_K - Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model file size&lt;/td&gt;
&lt;td&gt;4.86 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Runtime context&lt;/td&gt;
&lt;td&gt;4,096 tokens, one slot&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CPU inference threads&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;This is the configuration in my logs, not a specification for every Surface Book 3.&lt;/p&gt;
&lt;p&gt;Ollama offloaded &lt;strong&gt;28 of 37 layers&lt;/strong&gt; to the GPU: 27 repeating layers and the output layer. The model buffers were &lt;strong&gt;3,598.37 MiB on the GPU&lt;/strong&gt; and &lt;strong&gt;1,379.26 MiB in CPU-mapped memory&lt;/strong&gt;. The context cache added another &lt;strong&gt;432 MiB on the GPU&lt;/strong&gt; and &lt;strong&gt;144 MiB on the CPU&lt;/strong&gt;, with additional compute buffers on top.&lt;/p&gt;
&lt;p&gt;The memory-fitting log explains why this was not a fully GPU-resident run. On October 7, the initial estimate for full offload was &lt;strong&gt;5,311 MiB&lt;/strong&gt;, against &lt;strong&gt;5,199 MiB free&lt;/strong&gt;, and the fitter also aimed to leave &lt;strong&gt;1,024 MiB&lt;/strong&gt; free. It reduced GPU offload to fit that target.&lt;/p&gt;
&lt;p&gt;The practical takeaway is that this setup uses the GPU, but it is a &lt;strong&gt;mixed CPU/GPU workload&lt;/strong&gt;. The model&amp;rsquo;s download size alone does not describe its runtime memory requirement.&lt;/p&gt;
&lt;h2 id="turning-the-surface-book-into-a-lan-ollama-server"&gt;Turning the Surface Book into a LAN Ollama Server&lt;/h2&gt;
&lt;p&gt;By default, Ollama listens on &lt;code&gt;127.0.0.1:11434&lt;/code&gt;. That address is only reachable from the Surface Book 3 itself.&lt;/p&gt;
&lt;p&gt;In my setup, the Ollama desktop app&amp;rsquo;s tray startup did not apply the &lt;code&gt;OLLAMA_HOST&lt;/code&gt; environment variable to its own server process. The app started its server on localhost, even though I had configured a network address. This behavior was specific to the version and configuration I tested; check the behavior of your installed Ollama release.&lt;/p&gt;
&lt;p&gt;I used a separate &lt;code&gt;ollama.exe serve&lt;/code&gt; process instead:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;I set the Windows user environment variable &lt;code&gt;OLLAMA_HOST&lt;/code&gt; to &lt;code&gt;0.0.0.0:11434&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;I disabled the Ollama tray-app shortcut in the Startup folder to prevent it from launching another server on localhost.&lt;/li&gt;
&lt;li&gt;I started &lt;code&gt;ollama.exe serve&lt;/code&gt; as a background process instead.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Because &lt;code&gt;0.0.0.0&lt;/code&gt; binds Ollama on all IPv4 interfaces, the firewall rule is essential. I created an inbound rule for TCP port &lt;strong&gt;11434&lt;/strong&gt; on the &lt;strong&gt;Private&lt;/strong&gt; network profile. My setup verified listeners on both IPv4 (&lt;code&gt;0.0.0.0:11434&lt;/code&gt;) and IPv6 (&lt;code&gt;[::]:11434&lt;/code&gt;).&lt;/p&gt;
&lt;p&gt;For a more restricted setup, bind Ollama to a specific private address if your installed server supports it, and scope the firewall rule&amp;rsquo;s remote address to the Surface Laptop. Otherwise, permit the port only on a trusted Private profile, never on Public, and do not forward it through your router. The default Ollama API has no authentication and this connection uses unencrypted HTTP.&lt;/p&gt;
&lt;h3 id="starting-the-server-automatically"&gt;Starting the Server Automatically&lt;/h3&gt;
&lt;p&gt;To keep the standalone server available after I sign in to Windows, I created a Scheduled Task that starts &lt;code&gt;ollama.exe serve&lt;/code&gt; with a hidden PowerShell window. With &lt;code&gt;OLLAMA_HOST&lt;/code&gt; saved as a user environment variable, the process inherits it when the task starts at logon.&lt;/p&gt;
&lt;p&gt;This was more reliable for my server setup than the tray app, which both ignored the desired listening address and had GPU-detection timeouts during initial attempts. I verified the task by stopping the Ollama processes, starting the task, and checking that the server was listening and the model was still available.&lt;/p&gt;
&lt;h3 id="testing-network-access"&gt;Testing Network Access&lt;/h3&gt;
&lt;p&gt;First, check that the Ollama server responds on the Surface Book:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-powershell" data-lang="powershell"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nb"&gt;Invoke-WebRequest&lt;/span&gt; &lt;span class="n"&gt;-Uri&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;http://localhost:11434/api/version&amp;#34;&lt;/span&gt; &lt;span class="n"&gt;-UseBasicParsing&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Then, from the Surface Laptop, check the tags endpoint using the Surface Book&amp;rsquo;s private address:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;http://&amp;lt;surface-book-private-ip&amp;gt;:11434/api/tags
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The response should list the models available on the Surface Book, including &lt;code&gt;qwen3:8b&lt;/code&gt;. I also tested a prompt through the Ollama API from the other device and confirmed that the model returned a response.&lt;/p&gt;
&lt;blockquote class="border-l-4 border-neutral-300 dark:border-neutral-600 pl-4 italic text-neutral-600 dark:text-neutral-400 my-6"&gt;
&lt;p&gt;&lt;strong&gt;Important:&lt;/strong&gt; Do not expose port 11434 directly to the internet. For access outside the home network, use a properly secured VPN or an authenticated, encrypted gateway.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id="using-qwen38b-in-vs-code-chat"&gt;Using Qwen3:8b in VS Code Chat&lt;/h2&gt;
&lt;p&gt;VS Code Chat and GitHub Copilot CLI are separate clients. Configuring one does not automatically configure the other.&lt;/p&gt;
&lt;p&gt;For VS Code Chat, the essential step is to add a model provider that points to the &lt;strong&gt;Surface Book&lt;/strong&gt;, not to &lt;code&gt;localhost&lt;/code&gt; on the Surface Laptop.&lt;/p&gt;
&lt;p&gt;VS Code&amp;rsquo;s current documentation recommends the official
for Ollama models. Its default discovery address is local to the client, so a two-device setup needs a remote endpoint configuration.&lt;/p&gt;
&lt;p&gt;An explicit alternative is VS Code&amp;rsquo;s documented &lt;strong&gt;Custom Endpoint&lt;/strong&gt; provider:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Open Chat&amp;rsquo;s model picker and select &lt;strong&gt;Manage Language Models&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Select &lt;strong&gt;Add Models&lt;/strong&gt;, then &lt;strong&gt;Custom Endpoint&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Give the provider a recognizable name, such as &lt;strong&gt;Surface Book Ollama&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Select the &lt;strong&gt;Chat Completions&lt;/strong&gt; API type.&lt;/li&gt;
&lt;li&gt;Configure the model ID as &lt;code&gt;qwen3:8b&lt;/code&gt; and the model request URL as &lt;code&gt;http://&amp;lt;surface-book-private-ip&amp;gt;:11434/v1/chat/completions&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Follow the provider&amp;rsquo;s authentication prompts. Default Ollama does not validate an API key; a placeholder, if required by the client, does not secure the server.&lt;/li&gt;
&lt;li&gt;Save the configuration and select the model in Chat.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Provider labels and configuration screens can differ between VS Code versions. Follow the
for the fields your installed version exposes.&lt;/p&gt;
&lt;p&gt;For agent use, the client also needs accurate tool-calling and token-limit metadata. Do not advertise a larger context window than Ollama actually allocates.&lt;/p&gt;
&lt;p&gt;Start with a simple chat request before trying a multi-step agent task. Selecting this model for Chat does &lt;strong&gt;not&lt;/strong&gt; mean that inline code completions or every other Copilot feature use it.&lt;/p&gt;
&lt;h2 id="using-qwen38b-with-github-copilot-cli"&gt;Using Qwen3:8b with GitHub Copilot CLI&lt;/h2&gt;
&lt;p&gt;GitHub Copilot CLI can connect to an external model provider, including Ollama. In this setup, it runs on the Surface Laptop and sends model requests to the Surface Book.&lt;/p&gt;
&lt;p&gt;GitHub Copilot CLI supports Bring Your Own Key (BYOK) to connect to an OpenAI-compatible endpoint such as the Ollama API. For my tested setup, I set the provider URL and model in PowerShell on the &lt;strong&gt;Surface Laptop&lt;/strong&gt; before starting the CLI:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-powershell" data-lang="powershell"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nv"&gt;$env:COPILOT_PROVIDER_BASE_URL&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;http://&amp;lt;surface-book-private-ip&amp;gt;:11434/v1&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nv"&gt;$env:COPILOT_MODEL&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;qwen3:8b&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;copilot&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Replace the placeholder with the Surface Book&amp;rsquo;s private address. The settings above apply to that PowerShell session. To make them persistent for future sessions, Windows&amp;rsquo; &lt;code&gt;setx&lt;/code&gt; command can save user environment variables; open a new terminal afterward so it picks up the updated values. Do not set a provider API key for the default unauthenticated local Ollama server.&lt;/p&gt;
&lt;p&gt;My Ollama logs show successful requests to &lt;code&gt;/v1/chat/completions&lt;/code&gt; with a tool definition, confirming that the endpoint returned a response for a tool-enabled request. The current integration guidance can change, so check
and &lt;code&gt;copilot help providers&lt;/code&gt; for the options supported by your installed versions.&lt;/p&gt;
&lt;p&gt;The CLI needs a model that supports both &lt;strong&gt;streaming&lt;/strong&gt; and &lt;strong&gt;tool calling&lt;/strong&gt;. Availability in Ollama alone is not proof that every agent workflow will work well. Test a small request first, then a narrowly scoped coding task in a disposable working copy.&lt;/p&gt;
&lt;p&gt;The CLI tools operate on the Surface Laptop&amp;rsquo;s workspace. Running the model on the Surface Book does not automatically move repository files or command execution to that device.&lt;/p&gt;
&lt;p&gt;For an optional offline CLI session, GitHub documents &lt;code&gt;COPILOT_OFFLINE=true&lt;/code&gt;. This prevents Copilot CLI from contacting GitHub&amp;rsquo;s servers, but the Surface Laptop must still be able to reach the Ollama server over the LAN. It is not a network sandbox for commands or tools, and it does not configure VS Code Chat.&lt;/p&gt;
&lt;h2 id="real-performance-from-my-ollama-logs"&gt;Real Performance from My Ollama Logs&lt;/h2&gt;
&lt;p&gt;Two completed requests on October 6 give a useful snapshot of performance:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Measurement&lt;/th&gt;
&lt;th&gt;Request 1&lt;/th&gt;
&lt;th&gt;Request 2&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Prompt tokens evaluated&lt;/td&gt;
&lt;td&gt;2,050&lt;/td&gt;
&lt;td&gt;2,046, with 4 prompt tokens already cached&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompt evaluation time&lt;/td&gt;
&lt;td&gt;14.43 seconds&lt;/td&gt;
&lt;td&gt;14.20 seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompt evaluation speed&lt;/td&gt;
&lt;td&gt;142.02 tokens/second&lt;/td&gt;
&lt;td&gt;144.07 tokens/second&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Generated tokens&lt;/td&gt;
&lt;td&gt;362&lt;/td&gt;
&lt;td&gt;387&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Generation time&lt;/td&gt;
&lt;td&gt;41.11 seconds&lt;/td&gt;
&lt;td&gt;43.11 seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Generation speed&lt;/td&gt;
&lt;td&gt;8.78 tokens/second&lt;/td&gt;
&lt;td&gt;8.95 tokens/second&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Combined prompt and generation time&lt;/td&gt;
&lt;td&gt;55.55 seconds&lt;/td&gt;
&lt;td&gt;57.32 seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HTTP request duration reported by Ollama&lt;/td&gt;
&lt;td&gt;About 66 seconds&lt;/td&gt;
&lt;td&gt;58.78 seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;Prompt evaluation and response generation are different measurements.&lt;/strong&gt; The roughly 142-144 tokens per second describe processing the retained input, not writing the answer. For response generation, these two requests achieved approximately &lt;strong&gt;9 tokens per second&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The runner startup messages show approximately &lt;strong&gt;8.36 seconds&lt;/strong&gt; on October 6 and &lt;strong&gt;8.87-8.88 seconds&lt;/strong&gt; on October 7. These are runner startup timings, not time-to-first-token measurements. The HTTP request duration includes overhead beyond the combined evaluation timings; these values should not be treated as interchangeable.&lt;/p&gt;
&lt;p&gt;Short-interval generation readings sometimes exceeded 10 tokens per second, but the completed-request averages are a more useful figure to report than the fastest brief interval.&lt;/p&gt;
&lt;p&gt;For these particular responses, generation alone took roughly 41-43 seconds. That gives a realistic sense of the wait involved, but two requests cannot establish performance across different prompt lengths, system loads, or context settings.&lt;/p&gt;
&lt;h2 id="the-main-limitation-input-context-was-truncated"&gt;The Main Limitation: Input Context Was Truncated&lt;/h2&gt;
&lt;p&gt;An 8B model is not an 8 GB memory requirement. Model weights, quantization, the context cache, and other runtime allocations all affect memory use.&lt;/p&gt;
&lt;p&gt;More importantly, my logs show that the runtime used a &lt;strong&gt;4,096-token context&lt;/strong&gt;, while the model metadata reported &lt;code&gt;n_ctx_train = 40960&lt;/code&gt;. The runtime explicitly warned that it was not using the model&amp;rsquo;s full context capacity.&lt;/p&gt;
&lt;p&gt;The incoming requests were much larger than that runtime allocation:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Request&lt;/th&gt;
&lt;th&gt;Input prompt tokens reported before truncation&lt;/th&gt;
&lt;th&gt;Retained prompt tokens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;October 6, first completed request&lt;/td&gt;
&lt;td&gt;46,962&lt;/td&gt;
&lt;td&gt;2,050&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;October 6, second completed request&lt;/td&gt;
&lt;td&gt;46,964&lt;/td&gt;
&lt;td&gt;2,050&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;October 7, failed request&lt;/td&gt;
&lt;td&gt;59,382&lt;/td&gt;
&lt;td&gt;2,050&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;For example, the log reported:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;truncating input prompt limit=2050 prompt=46962 keep=4 new=2050
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That is a significant loss of input context. &lt;strong&gt;The successful timings above are for the retained prompt, not for processing all 47,000 input tokens.&lt;/strong&gt; An HTTP 200 response confirms that a request completed; it does not prove that the model received all the intended repository context or produced a correct answer.&lt;/p&gt;
&lt;p&gt;The excerpt also contains an HTTP &lt;strong&gt;404&lt;/strong&gt; on October 6 and an HTTP &lt;strong&gt;500&lt;/strong&gt;, followed by task cancellation, on October 7. It does not include enough error detail to establish their causes. I would not attribute either failure solely to context size based on this excerpt.&lt;/p&gt;
&lt;p&gt;Coding agents can send much more context than a simple chat prompt. Ollama&amp;rsquo;s Copilot CLI guide recommends at least &lt;strong&gt;64k tokens&lt;/strong&gt;, while GitHub recommends at least &lt;strong&gt;128k&lt;/strong&gt; for best results. These are recommendations, not a reason to configure an unsupported context size or exceed the Surface Book&amp;rsquo;s available memory.&lt;/p&gt;
&lt;p&gt;For this particular run, neither recommendation was met, and both exceed the 40,960-token training-context value reported in the model metadata. Simply changing a setting to 64k or 128k is not a verified fix. Any extended-context configuration needs model-specific support and separate validation.&lt;/p&gt;
&lt;p&gt;The practical starting point is to reduce the context sent by the client: use smaller excerpts, focused tasks, and shorter conversations. If increasing Ollama&amp;rsquo;s context allocation within the model&amp;rsquo;s supported configuration, check memory use, CPU/GPU allocation, response quality, and whether truncation warnings remain. The performance figures here do not establish how a larger context will behave.&lt;/p&gt;
&lt;p&gt;Useful starting tasks include explaining a small script, discussing a focused code change, or drafting documentation from a short excerpt. Treat these as workloads to evaluate, not as a promise that Qwen3:8b will match larger hosted coding models.&lt;/p&gt;
&lt;p&gt;Always review generated code and commands before using them. A local model can make mistakes just like a hosted model.&lt;/p&gt;
&lt;h2 id="what-stays-local-and-what-does-not"&gt;What Stays Local, and What Does Not&lt;/h2&gt;
&lt;p&gt;With this local model selected, inference takes place on the Surface Book. Prompts and code context sent for that inference travel from the Surface Laptop to Ollama over the private network.&lt;/p&gt;
&lt;p&gt;That is different from guaranteeing that no software contacts external services. Model downloads, editor extensions, telemetry, connected tools, and other selected providers may still use the internet.&lt;/p&gt;
&lt;p&gt;GitHub documents &lt;code&gt;COPILOT_OFFLINE=true&lt;/code&gt; for preventing Copilot CLI from contacting GitHub&amp;rsquo;s servers. If you use it, the configured provider must still be reachable: here that is the Surface Book on the private network. It does not block arbitrary network access by commands or tools the agent runs, and it does not configure VS Code Chat.&lt;/p&gt;
&lt;p&gt;For sensitive or business data, check organizational policy as well as the model endpoint. Keeping inference on your own hardware is only one part of protecting the workflow.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;My setup gives each Surface a clear role: the &lt;strong&gt;Surface Book 3&lt;/strong&gt; runs &lt;strong&gt;Ollama and Qwen3:8b&lt;/strong&gt;, while the &lt;strong&gt;Surface Laptop 7 for Business&lt;/strong&gt; provides the &lt;strong&gt;VS Code Chat and GitHub Copilot CLI&lt;/strong&gt; interfaces.&lt;/p&gt;
&lt;p&gt;The key is not simply installing a model. It is connecting both clients to the correct host, keeping that endpoint private, and understanding the limits of the model and hardware.&lt;/p&gt;
&lt;p&gt;My logs show approximately &lt;strong&gt;9 generated tokens per second&lt;/strong&gt; with partial GPU offload. They also show the more important constraint for coding-agent work: large prompts were heavily truncated. This makes the setup an interesting local AI experiment, but these measurements do not demonstrate reliable full-repository reasoning.&lt;/p&gt;
&lt;p&gt;It is a practical way to explore local AI with Surface hardware without requiring the daily-driver laptop to host the model itself.&lt;/p&gt;
&lt;h2 id="references"&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;</description></item></channel></rss>