AI Research Engineer, Inference

Op Recruiting ·www.oprecruiting.com

Location Chicago, IL, United States
Salary USD 250,000 - 300,000 / year
Type Full time
Level Mid
Source Shazamme
Trading Firm Accepting Candidates
Apply direct

Job Title: AI Research Engineer – Deep Learning Inference

Location: New York, NY | London, UK

About the Opportunity

Join an industry-leading global quantitative technology organization at the forefront of machine learning innovation. We are seeking a High-Performance AI Research Engineer to drive speed and scalability across our real-time predictive modeling infrastructure. In this role, you will bridge the gap between machine learning research and hardware acceleration, building low-latency inference systems that process continuous global data streams to drive core business decisions.

Responsibilities

  • Drive performance optimizations across all aspects of large-scale model execution, including custom kernel creation, data streaming pipelines, and novel hardware integration.

  • Partner directly with machine learning researchers to co-design neural architectures optimized for extreme real-time execution.

  • Author low-level primitives and custom operations to extract maximum compute throughput from modern hardware architectures.

  • Evaluate, benchmark, and deploy novel hardware acceleration tech, ranging from off-the-shelf accelerators to custom specialized silicon.

  • Formulate and execute engineering initiatives that address complex, non-obvious bottlenecks in ultra-low-latency deep learning inference.

Requirements (Must-Have)

  • At least two years of hands-on experience engineering production-grade deep learning systems within any complex domain (such as robotics, computer vision, audio, NLP, physics, or recommender platforms).

  • Strong lower-level engineering foundation, including experience writing custom compute kernels (e.g., CUDA, Triton, Pallas, or CuTe DSLs).

  • Proficiency with framework compilation internals and low-level runtime environments (PyTorch, JAX, XLA, or CUDA Graphs).

  • Practical exposure to hardware acceleration tech, such as FPGAs, ASICs, or specialized AI processors.

  • Proven ability to adapt algorithms and technical concepts across different domain applications.

Preferred Qualifications

  • Experience optimizing or serving Large Language Models (LLMs) and foundation architectures.

  • Note: Prior background in quantitative finance or trading is explicitly NOT required.

Compensation & Benefits

  • Highly competitive base salary, performance-based bonus incentive, and premium health/wellness benefits package. Equal Opportunity Employer.

Frequently asked questions

Who is hiring for the AI Research Engineer, Inference role?
Op Recruiting is hiring for the AI Research Engineer, Inference position, a Shazamme client. Apply directly on the employer's career site.
Where is the AI Research Engineer, Inference job located?
The AI Research Engineer, Inference role with Op Recruiting is based in Chicago, IL, US.
What does the AI Research Engineer, Inference role pay?
Op Recruiting lists the AI Research Engineer, Inference role at USD 250,000–300,000 per year.
Is the AI Research Engineer, Inference role full-time or contract?
This is a full time position at Op Recruiting.
What experience level is the AI Research Engineer, Inference role?
The AI Research Engineer, Inference position is aimed at mid-level candidates.
How do I apply for the AI Research Engineer, Inference role at Op Recruiting?
Apply directly on Op Recruiting's career page via the Apply button on this listing. ZammeJobs links straight through to the employer's ATS — no third-party form, no resume database.
Apply direct