TECH Signal 194
Morph (YC S23) Is Hiring Member of Technical Stuff
The role focuses on closing the gap between theoretical and actual hardware performance for model inference, which directly impacts the cost and speed of serving large models. It signals that specialized inference optimization—from kernels to routing—is a critical bottleneck for code-generation AI companies.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
The position requires deep expertise across the inference stack, including GPU performance, memory bandwidth, and distributed execution.
Responsibilities include tracing latency from the API layer to individual kernels and optimizing batching, scheduling, and quantization.
The company builds its own custom inference stack and speculative-decoding models rather than relying on off-the-shelf serving frameworks.
THE CLUSTER
↗