B

Systems & Research Engineer

Big Wave Digital San Francisco Bay Area
Remote
Apply
AI Summary

Design and optimize production AI systems, focusing on inference, model serving, and performance engineering for freight and supply chain applications. Profile bottlenecks in GPU, CPU, memory, and network resources to improve throughput, latency, and cost efficiency. Demonstrate exceptional technical depth by designing and executing rigorous benchmarks and experiments to drive engineering decisions.

Key Highlights
100% remote role across the USA with fully asynchronous operations
Total compensation of USD $220K–$300K plus equity
High-bar role requiring evidence of exceptional work in AI systems and performance engineering
Key Responsibilities
Profile production AI systems to identify GPU, CPU, memory, or network bottlenecks
Benchmark serving frameworks such as vLLM and SGLang to evaluate performance
Investigate inference throughput and latency to optimize workloads
Evaluate open-source versus closed-source models for quality and economics
Build evaluation infrastructure and understand model behavior under real production traffic
Turn research findings into production architecture decisions
Technical Skills Required
Model Serving Inference Optimization Performance Engineering
Benefits & Perks
100% Remote Work
Equity Compensation
Optional San Francisco Office Access
Nice to Have
Experience working with modern coding agents

Job Description



“I’ve seen things you people wouldn’t believe.”

Blade Runner


Now we’d like to see what you’ve built.


We’re recruiting a Systems & Research Engineer for a fast-growing US applied AI company building production systems for freight and global supply chains.


Founded by engineers from MIT and Stanford, the company raised $6M in seed funding in early2026 and is already operating AI products at serious real-world scale.


One of its core fraud and identity platforms now screens approximately 3,000 drivers every day.

The team is small.

The ambition isn’t.

And the engineering bar is deliberately high.


What You’ll Actually Do


This is not another generic AI Engineer role.

You’ll sit at the intersection of:


AI systems.

Research.

Inference.

Model serving.

Performance engineering.

Distributed systems.

Evaluation.

Your job is to understand how production AI systems actually behave.

Where is the bottleneck?

GPU?

CPU?

Memory?

Network?

Serving architecture?

Concurrency?

Model choice?


You’ll form hypotheses, build benchmarks, test alternatives and use the results to make real engineering decisions.


A recent example involved benchmarking speech-to-text approaches and building a hybrid open-source and production system that outperformed vendor alternatives on both quality and economics.

That’s the level of problem we’re talking about.


We Want the Experiment, Not Just the Percentage

A resume saying:

“Reduced inference latency by 37%.”

isn’t enough.

We want to know:

What was the baseline?

What did you think was happening?

How did you test it?

What alternatives did you benchmark?

What did the data show?

And most importantly:

What engineering decision changed because of the experiment?


You need to be able to walk us through at least one serious benchmark or experiment you personally designed and ran involving areas such as:

Model serving

Inference

Evaluation

Retrieval

Speech systems

Agent infrastructure

The strongest candidates think like researchers but ship like engineers.

The role specifically requires performance engineering on AI/model systems rather than distributed systems with no model in the loop.


The Kind of Work You Could Be Doing


Profiling production AI systems and identifying GPU, CPU, memory or network bottlenecks.

Benchmarking serving frameworks such as vLLM, SGLang and alternative architectures.

Investigating inference throughput and latency.

Evaluating open-source versus closed-source models.

Optimizing workloads across cost, quality and concurrency.

Building evaluation infrastructure.

Understanding model behavior under real production traffic.

Reasoning about distributed systems where there is genuinely a model in the loop.

Turning research findings into production architecture decisions.

And occasionally proving that everyone’s first assumption was wrong.


Who Could Be Right?

Your current title might be:

Research Engineer

ML Systems Engineer

AI Infrastructure Engineer

Inference Engineer

Performance Engineer

ML Platform Engineer

Systems Engineer


Experience inside sophisticated ML infrastructure environments is highly relevant.

Think engineering problems similar to those encountered at Google DeepMind, Meta, Stripe, Amazon AGI, Anthropic, Together AI, Fireworks AI, Baseten or similarly strong AI organizations.


But pedigree alone won’t get you through.

We’re looking for evidence of exceptional technical work.

The strongest candidates can explain their experience like this:

Here was the hypothesis.

Here was the baseline.

Here was the benchmark.

Here’s what we discovered.

Here’s what we changed.

What This Role Is NOT

This is not DevOps.

It isn’t frontend.

It isn’t conventional full-stack development.

It isn’t Web3.

And pure distributed-systems experience, however impressive, isn’t enough if there has been no meaningful AI or model component.


There needs to be a model in the loop.

AI Coding Agents

This team uses coding agents heavily.

Every day.

The philosophy is simple: excellent engineers should increasingly spend their time deciding what should be built, how it should work and whether the result is correct, rather than manually producing every line.

Your ability to work effectively with modern coding agents will be assessed during the interview process.


Remote Really Means Remote


This role is 100% remote across the USA.

New York.

San Francisco.

Austin.

Seattle.

Boston.

Miami.

Denver.

Wherever you do your best work.

The company operates asynchronously, with no fixed working-hour or timezone-overlap requirement. There is a San Francisco office available if you want it, but attendance is not required.


Compensation


USD $220K–$300K Total Compensation + Equity

USD $300K is the ceiling.

This is high-bar, opportunistic hiring rather than a volume recruitment campaign. The company is building an ongoing pipeline and wants exceptional engineers, not simply more engineers.


The Bar


One of the founders has a very simple hiring philosophy:

He wants engineers joining the company who make the existing engineering team better.

That means the bar is high.

Deliberately.


If you’ve done genuinely exceptional work around AI systems, inference, model serving, evaluation or performance engineering, we want to hear from you.

And when you apply, don’t just tell us what you built.

Tell us about the experiment.


Systems & Research Engineer | Applied AI | USD $220K–$300K + Equity

100% Remote Across the USA | Fully Async | San Francisco Office Optional

Model Serving | Inference | ML Systems | Performance Engineering


Similar Jobs

Explore other opportunities that match your interests

Site Reliability Engineer

Programming
1w ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

wordbricks

San Francisco Bay Area

Senior Applied AI Research Engineer

Programming
2w ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Hightouch

San Francisco Bay Area

Data Engineering - Frontier AI Lab

Programming
3w ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

this is growth

San Francisco Bay Area

Subscribe our newsletter

New Things Will Always Update Regularly