The University of Texas at Austin

UT Austin M.S. Computer Science thesis · Ongoing research

Small Language Models
for X-Ray Alignment

We’re training a small language model to center X-ray diffraction spots in a simulator. The goal is to develop models that labs can fine-tune for their equipment and run on their own hardware.

David Katz1 · Elias Stengel-Eskin1 · Zhantao Chen2,3

What does a Bragg peak tell us?

The atoms in a crystal form a repeating pattern. When X-rays scatter from those atoms, the waves add together at certain angles and produce a strong signal called a Bragg peak. On a detector, a reflection from a single crystal can appear as a bright spot.

Where the peaks appear tells scientists about the spacing and orientation of the crystal lattice. How bright they are carries information about the atoms and their arrangement. Together, these measurements help scientists work out a material’s structure and track how it changes during an experiment.

Background: X-ray diffraction at Diamond Light Source.

Why alignment matters

Before measuring a selected reflection, scientists need to align the instrument so they can record it reliably. That often means adjusting motors, checking the detector, and repeating. Automating this process could save time and let researchers spend more of the experiment on the material itself.

The task in this study

We focus on one part of alignment: bringing a selected spot to the center of a simulated detector. Computer vision tracks the spot, and the model proposes adjustments to two motor axes, delta and chi. Centering prepares the reflection for measurement. Solving the crystal structure still requires more diffraction data and analysis.

Schematic 128 by 128 detector with an off-center spot and horizontal and vertical offsets from detector center.
Illustrative detector geometry, not a measured frame or trajectory. The task controls delta and chi.

Earlier work on AI-assisted X-ray experiments

Chen and colleagues demonstrated an AI system for X-ray experiments in simulation and at a synchrotron. This thesis focuses on training and evaluating a small model for one task within that kind of experiment: Bragg-spot centering.

Two-panel figure from Chen and colleagues showing a six-circle diffractometer and an agentic workflow for finding and optimizing primary and secondary reflections.
Prior work: agentic AI-assisted X-ray alignment. Reproduced without modification from Zhantao Chen et al., “An agentic artificially intelligent X-ray scientist,” Nature Machine Intelligence 8, 1075–1086 (2026), doi:10.1038/s42256-026-01261-5, under CC BY-NC-ND 4.0.

We start with one task that we can train and test directly.

Using the same simulator and evaluation cases lets us compare the language model with a small numerical controller and examine where each succeeds or fails.

Prior paper
Broader agentic workflow with virtual and physical demonstrations.
This thesis
Simulated relative delta/chi Bragg-spot centering only.
Long-term direction
Combine models trained for individual tasks into a larger experimental workflow.

How the controller works

Computer vision measures the spot in each 128 × 128 detector image. The language model receives those measurements as text, along with calibration information.

Conceptual control loop from simulator image and motor state through computer vision, Gemma, Python response checks, Python validation and authorization, and back to simulator execution.
Gemma 4 E4B-it proposes a relative motor adjustment. Python checks the request before the simulator executes it. The compact MLP baseline receives numerical inputs and uses the same movement checks. We record the model’s proposed actions and explanations so they can be checked against the measurements. These records do not give access to the model’s internal reasoning.

Observe

Computer vision locates the selected spot and tracks its position after each move.

Propose

Gemma proposes the next delta/chi adjustment. We compare the model before training, after supervised fine-tuning, and after additional reinforcement fine-tuning.

Check

Python checks the request. If it passes, the simulator applies the move and returns the next measurement.

Simulated task
Co3Sn2S2 · 128 × 128 detector · success within 0.5 pixel · at most 20 attempts
Controller input
32 fields describing current measurements and recent feedback, passed as text to Gemma and as numbers to the MLP
Learned controllers
Gemma 4 E4B-it with rank-16 LoRA for supervised fine-tuning (SFT), followed by reinforcement fine-tuning (RFT)
Numerical baseline
6,402 parameters · two 64-unit ReLU hidden layers · two relative-motor outputs

Centering a spot in three adjustments

Before and after simulated detector images. The tracked spot begins offset from detector center and is centered after three adjustments.
Before and after views from the same recorded clean SFT episode. Centroid error decreased from 5.101 pixels at reset to 0.242 pixels after three accepted adjustments, crossing the 0.5-pixel success threshold. Images are static re-renders from archived simulator states using a shared linear intensity scale; this one episode does not estimate aggregate performance.

Development results

We tested the controllers on the same 940 development cases from 43 initialization groups, with three training seeds per trained controller. Fine-tuning enabled Gemma to center spots in the simulator. The compact MLP also performed well.

Figure A · clean condition Every trained seed reached 897/940 by 20 actions.
Clean development success within 20 actions: zero-shot Gemma 37 of 940, all already centered at reset; MLP, SFT Gemma, and SFT plus RFT each 897 of 940.
Clean execution on 940 paired development cases. All 37 zero-shot successes were already centered at reset, so zero-shot Gemma made no successful off-center correction. Shared final outcomes do not imply identical trajectories or an unavoidable ceiling.
Figure B · mixed-fault condition The controllers differed more when given fewer actions.
Line chart of cumulative mixed-fault development success at 5, 10, and 20 actions. MLP, SFT Gemma, and SFT plus RFT converge near 910 to 913 successes by 20 actions, with larger differences at earlier budgets.
Cumulative success on the same 940 mixed-fault development cases. Lines are three-seed means; shaded bands show the observed seed minimum-to-maximum range, not confidence intervals. Success@20 ranged from 910–913 for MLP, 911–913 for SFT, and 910–913 for SFT+RFT. Clean and mixed-fault checkpoints were trained separately, so their difference is not a controlled causal estimate of fault effects.

Fine-tuning made the task work for Gemma. Before training, the model made no successful corrections from an off-center starting point. After supervised fine-tuning, it could center spots through repeated measurements and adjustments.

The MLP was a strong baseline. Its success rate was similar to trained Gemma within 20 actions, and higher at five actions under mixed faults. Similar rates here do not establish statistical equivalence.

RFT did not improve success within 20 actions. Additional reinforcement fine-tuning sometimes reduced success when fewer actions were allowed. The measured results did not show a clear Gemma advantage over the MLP.

Why run a model locally?

We want labs to be able to fine-tune a model for their equipment and run it themselves. With downloadable weights and suitable hardware, a small model can use the lab’s measurements and calibration instructions to propose actions that software checks before execution.

Keep experimental data in the lab

Measurements, unpublished sample details, and instrument logs can stay on the lab’s systems when inference and logging are configured locally. The lab controls access and decides what to keep.

Operate without a cloud connection

Once the model and its software are installed, it can run without a cloud API key or internet connection. This could suit an air-gapped lab network, provided the software, logging, and updates are also set up to work offline.

Fine-tune it and keep the version you tested

A lab can train a model on examples of its task, keep the resulting weights and adapters, and choose when to update them. It can also return to an earlier version if needed, subject to the model’s license. Training can take place on an HPC system before the model is moved to lab hardware.

Avoid sending every request over the internet

A local model does not need to wait for a response from a cloud service at each step. That removes network travel time and dependence on the service’s uptime and rate limits. The model still needs time to compute, so actual response times have to be measured on the chosen hardware.

Where cloud APIs fit

OpenAI and Claude APIs give labs access to capable models without running the hardware themselves. The tradeoff is needing credentials, a network connection, and an external service for each request. Providers offer data controls, but execution still happens outside the lab. Running locally gives the lab more control and makes it responsible for maintenance and security.

Our training and evaluation ran on TACC Vista. Testing on lab hardware, measuring response times, and setting up an air-gapped system are possible next steps.

Further reading: Gemma fine-tuning, running Transformers offline, and hosted API data controls.

Training and evaluation on Vista

We used Vista CPU nodes to run simulations, generate training examples with a deterministic controller, prepare data, and check outputs. Grace Hopper GPUs ran Gemma fine-tuning and inference. Slurm job arrays let us evaluate multiple controllers, conditions, and seeds in parallel.

This work used the Vista system at the Texas Advanced Computing Center.

What we want to test next

So far, we have evaluated centering in simulation, with clean motor execution and modeled motor faults. Future work could include:

  1. Test transfer across simulator and calibration changes.
  2. Test unfamiliar starting conditions and study why spot tracking fails.
  3. Measure hardware requirements and response times on a local lab system.
  4. Test operation on isolated networks and validate the controller on physical instruments.
  5. Study how models trained for individual tasks could work together on a larger alignment procedure.

People

David Katz

M.S. Computer Science researcher

Department of Computer Science
The University of Texas at Austin

Elias Stengel-Eskin

Assistant Professor, Computer Science

Department of Computer Science
The University of Texas at Austin

Zhantao Chen

Assistant Professor, Mechanical Engineering

Walker Department of Mechanical Engineering
Oden Institute for Computational Engineering and Sciences

David Katz

M.S. Computer Science · The University of Texas at Austin

davidkatz@utexas.edu