# PSSA: Non-Transformer LLM Written in Rust from Scratch

> Published 2026-09-30 · https://www.promptzone.com/arlo_suzuki/pssa-non-transformer-llm-written-in-rust-from-scratch-3kin

PSSA is a non-transformer language model written from scratch in Rust. The project reached the front page of Hacker News, where the discussion thread recorded 57 points and 14 comments.

The repository at https://github.com/Sparticle62ops/pssa contains the full implementation without reliance on existing transformer frameworks.

## What It Is and How It Works

PSSA replaces the standard attention mechanism with an alternative architecture developed entirely in Rust. The codebase implements tokenization, forward passes, and training loops without calling into PyTorch or TensorFlow.

Developers can inspect every matrix operation and gradient step directly in the source files. No external model weights or pretrained checkpoints are included.

## Available Specs and Benchmarks

The project page lists no parameter counts, training FLOPs, or inference latency figures. Early comments on Hacker News note the absence of published benchmarks against transformer baselines of similar size.

Without reported numbers, direct speed or accuracy comparisons remain unavailable.

## How to Try It

Clone the repository and build with the standard Rust toolchain:

```shell
git clone https://github.com/Sparticle62ops/pssa
cd pssa
cargo build --release
```

Run the binary on a small text corpus to verify training and inference loops. The README supplies the minimal configuration file format.

## Pros and Cons

- Pros
  - Complete control over every line of model code
  - No dependency on large ML frameworks reduces attack surface
  - Rust memory safety guarantees apply to the training loop

- Cons
  - No published performance numbers or scaling results
  - Limited community tooling compared with Hugging Face ecosystems
  - Training throughput on consumer GPUs remains unmeasured

## Alternatives and Comparisons

Standard transformer models such as Llama 3 and Mistral rely on attention layers and benefit from mature CUDA kernels. PSSA trades that ecosystem for a clean Rust implementation.

| Feature              | PSSA                  | Llama 3 / Mistral     |
|----------------------|-----------------------|-----------------------|
| Architecture         | Non-transformer       | Transformer           |
| Language             | Rust                    | Python + CUDA         |
| Published benchmarks | None                    | Extensive             |
| Framework dependencies | None                  | PyTorch / vLLM        |

## Who Should Use This

Researchers studying novel architectures benefit from the transparent Rust implementation. Teams needing immediate production throughput or benchmarked accuracy should continue with established transformer stacks.

Developers comfortable writing custom training loops in systems languages will find the project accessible.

## Bottom Line

PSSA demonstrates that a complete language model can be implemented outside the transformer paradigm and outside Python, yet it still lacks the measurements required for practical adoption decisions.

The project supplies a clean starting point for anyone willing to add those measurements themselves.