aijobs.net

Inference Systems Performance Architect

San Jose, California, United States

USD 245K-325K Senior-level Full Time

Apply Save
Found 20h ago
Tasks
Perks/Benefits
Skills/Tech-stack

Benchmarking | Bottleneck analysis | Capacity Planning | Continuous batching | Distributed Systems | KV cache | LLM Inference | Language Model | Large Language Model | Machine Learning | Performance Engineering | Profiling | Prompt Caching | Service Level | Service-Level Objectives | Simulation Modeling | Tail Latency | Workload modeling

Education

N/A

Roles

AI | AI Performance Architect | Architect | Performance Architect

Regions

North America

Countries

Costa Rica | United States

States

California, US

Cities

San Jose, California, US

Apply Save
Language: en Views: 1 Clicks: 0 Saves: 0

Related jobs