Chainlink Labs 需要 Software Engineer, Data Growth 人员 在哥伦比亚 | LinkedIn
关于 Chainlink Chainlink 是行业标准的预言机平台,将资本市场带入链上,并为大多数去中心化金融(DeFi)提供动力。Chainlink 堆栈提供了必要的数据、互...
加载中...
Runtime Inference Acceleration Engineer Own end-to-end performance for one or more model families on our managed inference runtime. We are 10–20× faster than hyperscaler inference today; the bar is to widen that gap on every release.
Team Runtime
Location San Francisco / Remote (US, EU)
Type Full-time
Stack CUDATritonROCm / HIPPyTorchC++17 / C++20NCCL / RCCLLinux perf, Nsight, rocprof
Apply by email
Hiring manager replies within 5 business days.
The team
The runtime team is the engineering side of the research lab. Six engineers today. We write the kernels, the scheduler, and the serving layer for managed inference. The runtime ships into customer production weekly.
Reports to the head of runtime. One on-call rotation shared across runtime, cluster SRE, and customer engineering.
The role
What you’ll do
01 Profile real customer workloads with Nsight, rocprof, and our internal tracer; turn the bottleneck into a paper-quality benchmark and a merged kernel.
02 Write fused, attention-aware kernels in CUDA and Triton (and ROCm/HIP for AMD MI300X) targeting H100, H200, B200, and B300.
03 Own one model family end-to-end: tokenizer through KV cache through speculative decoding, on every supported accelerator.
04 Design and ship the next round of long-context primitives — paged KV, ring attention, sliding-window cache eviction.
05 Co-author at least one external write-up per quarter: a blog post, a paper, or a kernel released to the open ecosystem.
06 Carry the runtime on-call rotation alongside cluster SRE and customer engineering — about one week per six.
The bar
What we’re looking for
Five-plus years of systems-level performance work, with at least two of those years deep in GPU kernels.
Strong CUDA and at least one of Triton, CUTLASS, or HIP. Comfortable reading PTX and SASS when the profile demands it.
Track record of measurable speed-ups landed in production or in widely-used open source — bring a PR or a perf graph.
Comfort with transformer internals: attention variants, KV cache shapes, quantization (FP8, INT4 / NF4), speculative decoding.
Empirical instincts. You design experiments, hold variables fixed, and write down the numbers before the next change.
Bonus
Nice to have, not required
Published kernels in vLLM, SGLang, TensorRT-LLM, FlashAttention, or similar.
Experience tuning collectives (NCCL / RCCL) and topology-aware scheduling.
AMD MI300X / MI325X experience — most candidates do not have this; we will weight it heavily.
Prior research experience or co-authorship on systems / ML papers.
Compensation
In writing, like everything else
We publish bands. We meet them. The number you see on the offer is the same number your future peers got at the same level. We do not negotiate; we level.
Base
$220,000 – $360,000 USD (US), depending on level (IC4 / IC5 / IC6).
Equity
Meaningful early-stage equity, refreshed on tenure milestones, not review cycles.
Notes
Pay bands are published internally and shared in the manager call. We do not negotiate; we level.
How to apply
One email is enough
Send a short note to [email protected] with the role title in the subject line. Include your CV or LinkedIn, one or two links to work you’re proud of, and a sentence on why this role specifically. Hiring managers reply within five business days, regardless of outcome.
Application
A hiring manager reads every email. Reply within five business days.
Manager call
30–45 minutes. Scope, role, mutual fit. We share the comp band on this call.
Technical loop
3–4 sessions on the same day. Real problems, no homework, no whiteboard riddles.
Offer
Same-week offer at the published band for your level. Start dates are flexible.
Equal opportunity
We hire on the work. Race, gender, age, nationality, religion, sexual orientation, disability, and veteran status do not factor into our decisions. We sponsor visas for senior roles in the US, UK, and EU — bring it up on the manager call.
Need an accommodation for the interview process? Mention it in your application or write [email protected].
Also open
Other roles you might consider
Research lab
Distributed Training Researcher
Lead a paper / quarter on multi-thousand-GPU pre-training. Co-appointment with a partner university available.
View role
Cluster & SRE
Cluster Site Reliability Engineer
Bring up B300 racks, drive InfiniBand fabric to spec, run capacity planning across seven regions. Pager included.
View role
Customer engineering
Customer Cluster Engineer
Embedded with a small portfolio of reserved-tier accounts. Distributed training perf, NCCL, kernel tuning.
View role
All open roles
One last thing
If this role isn’t quite right but you’d be a fit at iframe.ai, write anyway.
Senior engineers and researchers can apply outside the listed roles. The bar is the same. The reply window is the same.
Apply: Inference Acceleration Engineer General application
Five-plus years of systems-level performance work, with at least two of those years deep in GPU kernels.
Strong CUDA and at least one of Triton, CUTLASS, or HIP. Comfortable reading PTX and SASS when the profile demands it.
Track record of measurable speed-ups landed in production or in widely-used open source — bring a PR or a perf graph.
Comfort with transformer internals: attention variants, KV cache shapes, quantization (FP8, INT4 / NF4), speculative decoding.
Empirical instincts. You design experiments, hold variables fixed, and write down the numbers before the next change.
Meaningful early-stage equity, refreshed on tenure milestones, not review cycles.
Pay bands are published internally and shared in the manager call. We do not negotiate; we level.
The runtime team is the engineering side of the research lab. Six engineers today. We write the kernels, the scheduler, and the serving layer for managed inference. The runtime ships into customer production weekly.
One email is enough
Send a short note to [email protected] with the role title in the subject line. Include your CV or LinkedIn, one or two links to work you’re proud of, and a sentence on why this role specifically. Hiring managers reply within five business days, regardless of outcome.
注册并登录后即可查看
关于 Chainlink Chainlink 是行业标准的预言机平台,将资本市场带入链上,并为大多数去中心化金融(DeFi)提供动力。Chainlink 堆栈提供了必要的数据、互...
Fulcrum技术解决方案公司正在寻找一名Cribl工程师,负责支持企业级数据管道环境的迁移、优化和管理。该职位将专注于将现有数据流迁移到Cribl平台,配置和维护Cribl基础设施,并支持网络安全数据摄入项目。
职位:软件工程师 - 电商(远程) 位置:远程(可在任何地方工作) 工作类型:全职 薪资:$180K - $250K/年 职位概述:我们正在为其中一位客户招聘一名软件工程师,新毕业生(Zara)以全职工作为基础。该职位涉及开发可扩展的后端系统,优化算法,并与跨职能团队合作交付高性能解决方案。您将参与支持全球平台实时用户体验的关键基础设施。
职位名称:全栈开发人员 经验:3年以上 工作地点:埃及 类型:全职 薪资:根据经验竞争 关于我们 我们是一家获得风投支持的技术初创公司,致力于在医疗领域开发人工智能驱动的平台...