talentyGo

Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group

📍 Bala Cynwyd, Pennsylvania, US0💼 Full-time🕐 5/29/2026
Apply now →

Start your Pro trial in 30 seconds (3 days free): you also get the AI match score with your resume.

Role overview

Susquehanna International Group is hiring for the Machine Learning Engineer- Inference Optimization | Experienced Hire role in Bala Cynwyd, Pennsylvania, US. It is full-time, Mid-level level, in the Tech sector. It was posted 5/29/2026.

On TalentyGo you can review this job and apply more effectively: Charlie prepares an ATS-optimized resume and a cover letter tailored to "Machine Learning Engineer- Inference Optimization | Experienced Hire" at Susquehanna International Group in about a minute. Before you apply, you can also check how well your profile fits, with a match score based on skills, experience, location and seniority.

Role
Machine Learning Engineer- Inference Optimization | Experienced Hire
Company
Susquehanna International Group
Location
Bala Cynwyd, Pennsylvania, US
Work mode
On-site
Employment
Full-time
Seniority
Mid-level
Sector
Tech
Posted
5/29/2026

Description

Overview We are looking for a Machine Learning Engineer focused on low-latency inference optimization to help build, tune, and productionize high-performance model serving systems. This role sits at the intersection of machine learning, systems engineering, and GPU performance. You will work on inference workloads where latency, throughput, reliability, and hardware efficiency all matter, and where a deep understanding of modern inference runtimes can meaningfully improve production outcomes. You will work closely with quantitative researchers and engineers to understand model structure, identify inference bottlenecks, and turn research ideas into efficient production systems. The work may involve other types of models, but focuses on transformer-style architectures, and structured inference workloads. You will evaluate and tune frameworks and related serving or compilation systems, while also reasoning about GPU execution, memory layout, batching strategies, precision tradeoffs, and end-to-end latency. What you'll do • Design, build, and optimize low-latency inference systems for production machine learning workloads. • Profile model inference pipelines across model execution, runtime configuration, batching, memory movement, serialization, networking, and I/O. • Evaluate, integrate, and tune inference runtime systems. • Improve latency, throughput, GPU utilization, for production inference workloads. • Build and support benchmarking and profiling tools to compare model variants, hardware targets, runtime configurations, and deployment strategies. • Debug performance issues involving GPU memory, compute saturation, kernel behavior, CPU/GPU coordination, data movement, and serving-layer overhead. • Help shape model and system design choices so that research models are efficient to deploy under real latency constraints. • Where necessary, collaborate with lower-level systems or GPU specialists on custom operators, kernel-level optimization, or hardware-specific performance work. What We're Looking For • Experience deploying, optimizing, or operating machine learning inference workloads in production or production-like environments. • Programming experience in Python, Java, C# etc. and at least one systems language such as C, C++, Rust, or Go • Solid understanding of modern ML frameworks such as PyTorch, including model execution, export, tracing, compilation, and performance profiling. • Ability to reason about latency, throughput, batching, memory use, GPU utilization, and reliability under real workloads. • Strong practical judgment around tradeoffs between model quality, latency, throughput, implementation complexity, and maintainability. Preferred Qualifications • Experience optimizing inference for latency-sensitive or high-throughput applications. • Experience with model optimization techniques such as quantization, pruning, distillation, operator fusion, graph lowering, custom operators, or model compilation. • Exposure to CUDA, Triton language, ROCm, PTX, CuTe, CUTLASS, FlashInfer, or similar low-level GPU programming tools. • Experience running inference workloads on Kubernetes or GPU clusters, including scheduling, autoscaling, observability, and resource management. • Background in mathematics, physics, computer science, engineering, statistics, quantitative finance, or another technical field. • Demonstrated ability to improve real-world inference performance beyond a baseline framework implementation. If you're a recruiting agency and want to partner with us, please reach out to [email protected]. Any resume or referral submitted in the absence of a signed agreement will not be eligible for an agency fee.

The market for this role in Bala Cynwyd

TalentyGo lists 5862 similar roles (2 in Bala Cynwyd), 24% remote. Charlie ranks them against your CV, each with a clear score.

Similar roles
5862
Remote
24%
New / 7d
107

Similar jobs

Frontend Engineer III, Moderation Enforcement
📍 Remote - United States · Remote · tech
Senior Software Engineer (AI)
📍 Tel Aviv · tech
Senior Privacy Engineer
📍 Remote · Remote · tech
Senior Software Engineer, Application Security
📍 San Francisco Bay Area · tech
Director, Engineering
📍 Tel Aviv · tech
Forward Deployed Engineer, Finance [Office of the CTO]
📍 Remote - USA · Remote · finance
See all similar jobs →
Apply now →

TalentyGo is an aggregator of job postings from public sources. Always verify information directly with the company. Applications go through the original company website; TalentyGo does not manage hiring processes.