Data Infrastructure Engineer
Alljoined · San Francisco
About this role
## About Alljoined Alljoined is creating a future where humans are fully understood and augmented by technology. Our work solves the communication bottleneck between humans and computers by decoding thoughts from the brain, entirely non-invasively. We apply deep learning research to large-scale EEG datasets to decode multimedia input, eventually moving to internal thought. We are state-of-the-art in capabilities and fully vertically integrated, and we’re building a general consumer interface to transform how we live. We’re actively growing our founding engineering team to build the underlying infrastructure that makes this ambitious future a reality. ## About the Role As a **Data Infrastructure Engineer**, you will build the backend and hardware architecture that enables high-quality, fast research. You’ll own the entire data lifecycle—from building pipelines that process massive multimodal datasets (video, audio, text, time-series) to provisioning and managing both cloud and bare-metal compute clusters used to train on it. You’ll power foundational model training by bridging physical neuro hardware and central repositories, working alongside world-class researchers to deliver a **high-throughput, low-latency pipeline directly to the GPUs**. ## You Might Be a Good Fit If You - Have **3+ years** of production software engineering experience with deep expertise in systems-level architecture and languages like **Python, Rust, C++, or Go**. - Have built and maintained **high-performance ETL pipelines** capable of processing, buffering, and storing **terabytes of daily unstructured data**. - Are comfortable architecting, provisioning, and maintaining **bare-metal local compute clusters**, storage servers, and **high-speed networking** for intensive ML workloads. - Can handle **continuous, highly concurrent data streams** from heterogeneous hardware peripherals **without data loss**. - Can work across **hybrid environments** to define storage topologies, manage databases (**TimescaleDB, ClickHouse**), and sync massive datasets between **on-premise edge servers** and the **cloud** (**AWS/GCP/Azure**). - Enjoy owning the full technical lifecycle of infrastructure—from optimizing low-level I/O bound operations to production deployment. ## Strong Candidates May Have - Deep understanding of modern ML frameworks (**PyTorch/TensorFlow**) and experience building datasets that maximize and saturate **GPU utilization**. - Experience managing networking for distributed GPU training (**InfiniBand, RoCE**) or optimizing **zero-copy networking** and **shared memory**. - Experience building infrastructure involving programmatic video processing (**FFmpeg, GStreamer, OpenCV**). ## Compensation Range **$140,000 – $180,000/year** Final compensation will be determined based on your specific skills and experience and may be outside this range. ## Benefits - Competitive equity compensation at a seed stage startup - Options for housing support - Visa sponsorship - **3% 401k matching** - Health insurance
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.