Programmable storage for AI data pipelines

Move less data. Keep AI compute fed.

Jitronix is a programmable-storage runtime and simulator for validating safe, bounded eBPF processing closer to NVMe storage—before committing to a controller integration.

Working PoC Verified results Design partners wanted

Verified PoC

AI dataset quality filter

Outputs match

Input dataset

640 KB

10,000 records

Returned

64 KB

1,000 matches

Device-to-host transfer

90% less

Measured in the Jitronix software PoC using synthetic fixed-size records and 10% selectivity. This is a data-reduction result—not a hardware latency or GPU-throughput claim.

Working software PoC

Byte-for-byte verified output

Seeking AI infrastructure design partners

Why now

AI compute is expensive. Moving irrelevant data makes it worse.

Jitronix explores a simple question: when data must be inspected or transformed before use, can a safe program near storage return only what the next stage needs?

Too many bytes move
AI pipelines often transfer full datasets even when only a fraction of the records are useful downstream.
Filtering lands on scarce compute
CPUs and accelerators spend time receiving and rejecting data that could have been reduced closer to storage.
The data path is the constraint
Faster compute increases pressure on storage, memory, and interconnects. Moving every byte does not scale for every workload.

Proof, not a promise

A working software PoC with a clear evidence boundary.

The current demonstrator executes an eBPF filter in a sandboxed storage-memory model, returns matching records, and independently compares every output byte with a host implementation.

640,000

input bytes

64,004

device-to-host bytes

1,000

matching records

Verified

against host output

1. Read in sandbox

10,000 synthetic records

2. Execute eBPF

Bounded quality-score filter

3. Return matches

90% less transfer

Conventional host path640 KB
Jitronix filtered path64 KB

The platform

A portable execution layer for programmable storage.

Jitronix gives infrastructure teams a controlled way to test whether moving a small, bounded function toward storage is worthwhile—before taking on firmware, silicon, or fleet risk.

  • Safe eBPF execution

    Validate programs and constrain memory access before logic reaches a storage target.

  • Reproducible workload evaluation

    Compare host and storage-side paths using the same input and byte-for-byte output checks.

  • Hardware-ready architecture

    A modular backend designed to progress from software simulation toward NVMe controller integrations.

NVMe storage infrastructure in a data centre

Development path

Simulator → workload evidence → controller target

Built toward the emerging NVMe Computational Programs model, with hardware validation as the next milestone.

Candidate workloads

Where storage-side execution could earn its place.

The right workload is selective, bounded, and cheaper to evaluate near the data than after moving every byte. We are validating that boundary with infrastructure and controller teams.

PoC implemented
AI dataset quality filtering
Select records by bounded metadata or quality criteria before transferring them into the host data path.
Design-partner candidate
Metadata pre-filtering
Reduce scans by returning only records or objects whose local metadata matches the next processing stage.
Design-partner candidate
Inline inspection
Evaluate bounded validation, classification, or policy checks where avoiding unnecessary transfer has measurable value.

Design partners

Bring the workload. We will test the thesis together.

We are looking for AI infrastructure, storage, and controller teams willing to pressure-test whether a real workload belongs closer to the data.

Technical fit review

Map where storage-side processing fits—and where it does not—in your data path.

Workload-specific PoC

Evaluate a bounded filter or transformation using representative data and success criteria.

Reproducible evidence

Compare correctness and data movement now, then define the hardware benchmark that matters.

Start with a 30-minute technical review.

No platform rollout and no hardware commitment. We will begin with your data path, selectivity, record shape, and current bottleneck.