llm
12 posts
-
Configure opencode to use TritonAI Developer API at UC San Diego
A step-by-step tutorial for adding UC San Diego's TritonAI Developer API as a provider in opencode, including how to request access, find available models, and test them.
-
Control your existing ChromeOS browser with any AI coding agent
A reliable architecture for safely connecting multiple AI coding agents to an existing ChromeOS browser session using the Playwright extension.
-
Deploy a ChatGPT-like LLM on Jetstream with llama.cpp, tested on g3.medium
This is a tested follow-up and updated standalone version of Deploy a ChatGPT-like LLM on Jetstream with llama.cpp. If you want the original September 2025 version for reference, see: Deploy a ChatGPT-like LLM on Jetstream with llama.cpp
-
Cosmosage: A specialized AI assistant for cosmology
Cosmosage is an AI assistant specialized in cosmology. It has been trained on thousands of cosmology papers and textbooks, starting from a large-language model base and subsequently fine-tuned on cosmology-specific data and synthetic Q&A pairs.
-
Deploy a ChatGPT-like LLM on Jetstream with llama.cpp
Tutorial last updated in December 2025 This is a crosspost of the official Jetstream documentation: Deploy a ChatGPT-like LLM service on Jetstream. I built a brand new version of that tutorial that swaps in llama.cpp for vLLM so we can run GGUF quantized models on Jetstream's GPUs without giving up speed or context length.
-
Deploy a 70B LLM to Jetstream
Deploying large language models on Jetstream is getting easier thanks to the official Jetstream LLM guide. Here I follow that walkthrough but scale the hardware and model so we can run something far more capable than the defaults.
-
Timing the unshelving of a Jetstream 70B LLM instance
Following the work documented in Deploy a 70B LLM to Jetstream, the Meta-Llama-3.1-70B-Instruct-GGUF deployment is now running on a g3.xl instance. The goal of this follow-up is to measure how long it takes to unshelve that virtual machine and bring the chat interface back online.
-
How to use AI chat assistants while traveling
Traveling often means unreliable or expensive internet access, especially on airplanes or in remote areas. Here are two practical ways to keep using AI chat assistants even when your connectivity is limited:
-
VS Code Copilot agent mode: A productivity showcase
Visual Studio Code, together with GitHub Copilot in agent mode, enables a more focused and efficient development workflow. Below are three practical examples of how this integration can help streamline your daily work, each illustrated with real screenshots:
-
Astrophysics Papers Daily Summaries Notebook with the Jetstream LLM Inference Service
Fetch today's astro-ph papers from arXiv, summarize abstracts with llm, and output a Markdown summary. This is an example of using the Jetstream Inference Service, notice that you need first to configure the llm package to access the Jetstream Inference Service via the API
-
Use Jetstream's DeepSeek R1 as a code assistant on JupyterAI
Thanks to Openrouter, there is now a way of using the Jetstream LLM inference system, in particular the powerful Deepseek R1 model, as a code and documentation assistant in JupyterLab via JupyterAI.
-
Deploy an LLM ChatGPT-like service on Jetstream
In this tutorial we will deploy a LLM Chat-GPT like service on a GPU node on Jetstream. However the same instructions can be used to deploy any other model available on the Hugging Face model hub.