High Performance Computing & AI
Practical notes on Python, JupyterHub, Kubernetes and AI for science — from the San Diego Supercomputer Center.
-
OpenStack Unshelver Demo
I recently vibe-coded a lightweight web application, using gpt-5-codex, that revives shelved OpenStack instances on demand. The stack is intentionally minimal: a FastHTML frontend, GitHub for authentication, and the OpenStack SDK orchestrated through a YAML configuration that lists the instances the team cares about.
-
Deploy a 70B LLM to Jetstream
Deploying large language models on Jetstream is getting easier thanks to the official Jetstream LLM guide. Here I follow that walkthrough but scale the hardware and model so we can run something far more capable than the defaults.
-
Timing the unshelving of a Jetstream 70B LLM instance
Following the work documented in Deploy a 70B LLM to Jetstream, the Meta-Llama-3.1-70B-Instruct-GGUF deployment is now running on a g3.xl instance. The goal of this follow-up is to measure how long it takes to unshelve that virtual machine and bring the chat interface back online.
-
Python for HPC
Python is often the first choice for prototyping research ideas, but scaling that prototype to thousands of cores and multi‑node workflows needs a different toolkit.
-
Execute Pegasus jobs on Expanse
Pegasus is a workflow management system that helps scientists and engineers execute complex computational workflows. Pegasus supports different data staging mechanisms, primarily sharedfs and condorio. The sharedfs mode is used when the head node and all worker nodes share a common file system, allowing jobs to directly access data. However, there is currently a bug that prevents sharedfs from working as expected. In contrast, the condorio mode is designed for environments where worker nodes do not share a file system, relying on HTCondor's built-in file transfer capabilities for all data I/O. Since Pegasus 5.0, condorio is the default. Our current setup on Expanse, utilizing HTCondor Annex, operates in condorio mode, leveraging Condor's efficient data transfer for distributed execution.
-
Use VS Code on Expanse
Using VS Code directly on Expanse, or other HPC systems, is generally not recommended due to the high resource usage of the IDE itself. These systems are optimized for computational tasks, not for running graphical applications on login nodes.
-
How to use AI chat assistants while traveling
Traveling often means unreliable or expensive internet access, especially on airplanes or in remote areas. Here are two practical ways to keep using AI chat assistants even when your connectivity is limited:
-
Voyager enters ACCESS allocations phase
As of June 1, 2025, SDSC's Voyager supercomputer has transitioned from its testbed phase to the ACCESS Allocations Phase. All research use now requires an official ACCESS allocation.
-
How to use GitHub Copilot for scientific computing
GitHub Copilot is rapidly transforming the landscape of scientific computing by streamlining code development and accelerating research workflows. This guide outlines how computational scientists can leverage Copilot, with a focus on professional and research-oriented use cases.
-
Python and AI coding summer camp in person in San Diego
I'm excited to announce that I'll be teaching a brand new Python and AI Coding Summer Camp in person in San Diego, hosted at the Italian School of San Diego (Kearny Mesa)!
-
VS Code Copilot agent mode: A productivity showcase
Visual Studio Code, together with GitHub Copilot in agent mode, enables a more focused and efficient development workflow. Below are three practical examples of how this integration can help streamline your daily work, each illustrated with real screenshots:
-
Astrophysics Papers Daily Summaries Notebook with the Jetstream LLM Inference Service
Fetch today's astro-ph papers from arXiv, summarize abstracts with llm, and output a Markdown summary. This is an example of using the Jetstream Inference Service, notice that you need first to configure the llm package to access the Jetstream Inference Service via the API
-
BK18 added to Panexp model suite and new WebSky radio source models in PySM 3.4.1
Following up on my previous post about Panexp model suite simulations, there have been several important updates: BK18 Simulation Now Available
-
Deploy a NFS server to share data between JupyterHub users on Jetstream
This is an updated version of my 2023 tutorial on deploying a NFS server to share data between JupyterHub users on Jetstream.
-
How to request High Performance Computing (HPC) resources
I am often asked how a scientist located in the US can access supercomputing resources. I decided to write a blog post with an overview of the options, consider that I'm writing this in April 2025, so please cross-check on the official websites.