Keep experimental data in the lab
Measurements, unpublished sample details, and instrument logs can stay on the lab’s systems when inference and logging are configured locally. The lab controls access and decides what to keep.
UT Austin M.S. Computer Science thesis · Ongoing research
We’re training a small language model to center X-ray diffraction spots in a simulator. The goal is to develop models that labs can fine-tune for their equipment and run on their own hardware.
The atoms in a crystal form a repeating pattern. When X-rays scatter from those atoms, the waves add together at certain angles and produce a strong signal called a Bragg peak. On a detector, a reflection from a single crystal can appear as a bright spot.
Where the peaks appear tells scientists about the spacing and orientation of the crystal lattice. How bright they are carries information about the atoms and their arrangement. Together, these measurements help scientists work out a material’s structure and track how it changes during an experiment.
Background: X-ray diffraction at Diamond Light Source.
Before measuring a selected reflection, scientists need to align the instrument so they can record it reliably. That often means adjusting motors, checking the detector, and repeating. Automating this process could save time and let researchers spend more of the experiment on the material itself.
We focus on one part of alignment: bringing a selected spot to the center of a simulated detector. Computer vision tracks the spot, and the model proposes adjustments to two motor axes, delta and chi. Centering prepares the reflection for measurement. Solving the crystal structure still requires more diffraction data and analysis.
Chen and colleagues demonstrated an AI system for X-ray experiments in simulation and at a synchrotron. This thesis focuses on training and evaluating a small model for one task within that kind of experiment: Bragg-spot centering.
We start with one task that we can train and test directly.
Using the same simulator and evaluation cases lets us compare the language model with a small numerical controller and examine where each succeeds or fails.
Computer vision measures the spot in each 128 × 128 detector image. The language model receives those measurements as text, along with calibration information.
Computer vision locates the selected spot and tracks its position after each move.
Gemma proposes the next delta/chi adjustment. We compare the model before training, after supervised fine-tuning, and after additional reinforcement fine-tuning.
Python checks the request. If it passes, the simulator applies the move and returns the next measurement.
We tested the controllers on the same 940 development cases from 43 initialization groups, with three training seeds per trained controller. Fine-tuning enabled Gemma to center spots in the simulator. The compact MLP also performed well.
Fine-tuning made the task work for Gemma. Before training, the model made no successful corrections from an off-center starting point. After supervised fine-tuning, it could center spots through repeated measurements and adjustments.
The MLP was a strong baseline. Its success rate was similar to trained Gemma within 20 actions, and higher at five actions under mixed faults. Similar rates here do not establish statistical equivalence.
RFT did not improve success within 20 actions. Additional reinforcement fine-tuning sometimes reduced success when fewer actions were allowed. The measured results did not show a clear Gemma advantage over the MLP.
We want labs to be able to fine-tune a model for their equipment and run it themselves. With downloadable weights and suitable hardware, a small model can use the lab’s measurements and calibration instructions to propose actions that software checks before execution.
Measurements, unpublished sample details, and instrument logs can stay on the lab’s systems when inference and logging are configured locally. The lab controls access and decides what to keep.
Once the model and its software are installed, it can run without a cloud API key or internet connection. This could suit an air-gapped lab network, provided the software, logging, and updates are also set up to work offline.
A lab can train a model on examples of its task, keep the resulting weights and adapters, and choose when to update them. It can also return to an earlier version if needed, subject to the model’s license. Training can take place on an HPC system before the model is moved to lab hardware.
A local model does not need to wait for a response from a cloud service at each step. That removes network travel time and dependence on the service’s uptime and rate limits. The model still needs time to compute, so actual response times have to be measured on the chosen hardware.
OpenAI and Claude APIs give labs access to capable models without running the hardware themselves. The tradeoff is needing credentials, a network connection, and an external service for each request. Providers offer data controls, but execution still happens outside the lab. Running locally gives the lab more control and makes it responsible for maintenance and security.
Our training and evaluation ran on TACC Vista. Testing on lab hardware, measuring response times, and setting up an air-gapped system are possible next steps.
Further reading: Gemma fine-tuning, running Transformers offline, and hosted API data controls.
Computing at TACC
We used Vista CPU nodes to run simulations, generate training examples with a deterministic controller, prepare data, and check outputs. Grace Hopper GPUs ran Gemma fine-tuning and inference. Slurm job arrays let us evaluate multiple controllers, conditions, and seeds in parallel.
This work used the Vista system at the Texas Advanced Computing Center.
So far, we have evaluated centering in simulation, with clean motor execution and modeled motor faults. Future work could include:
M.S. Computer Science researcher
Assistant Professor, Computer Science
Assistant Professor, Mechanical Engineering
Contact
M.S. Computer Science · The University of Texas at Austin