Forward Deployed Engineer
Role Description
About Tensormesh
Building the caching layer for AI production.
Tensormesh was founded by the creators of LMCache to bring production-grade context caching to AI teams. We help developers, enterprises, and industry-leading partners eliminate redundant prefill and KV-cache recompute, lower inference costs, and scale workloads that reuse context across requests. Our product runs on the Tensormesh platform or deploys directly into customers' own inference clusters and environments, cutting inference cost and latency by making KV cache a ubiquitous layer across every inference stack.
About the role
As a Forward Deployed Engineer, you are the last mile between what our product engineering team builds and what the customer actually runs in production. You take our technology (e.g. LMCache and more across our inference-optimization stack), stand it up in the customer's real environment, benchmark it, prove the value, and feed the hard problems back to engineering. You'll own deployments end-to-end and turn what you learn in the field into repeatable playbooks the whole FDE team can use. This is a high-ownership, customer-facing engineering role for someone equally comfortable with deep technical work and direct customer engagement.
What you'll do
- Deploy and tune our inference-optimization stack in customer environments: Kubernetes/Helm, sidecar, standalone, egress-restricted, and air-gapped setups.
- Design and run reproducible benchmarks that demonstrate KV-cache reuse and performance gains on the customer's own models and hardware.
- Diagnose performance and correctness issues in the field (e.g. eviction behavior, cache-hit ratios, TTFT/throughput regressions) and drive them to root cause with engineering.
- Build and standardize the FDE apparatus: validated deployment blueprints, feature-showcase flows, and benchmark/debug tooling, so the same input produces the same result for every customer and across your team members.
- Act as the technical voice of the customer: surface recurring blockers, feedback patterns, and required product actions back to engineering and leadership.
- Partner with sales/solutions during evaluations, POCs, and contract setups (including remote-access and security requirements)
What you'll bring
- Strong software / systems engineering background; comfortable in Python and Linux, reading and modifying production code.
- Hands-on experience deploying, operating, and debugging workloads on Kubernetes (Helm, CRDs, storage, add-ons, networking) in real customer or production settings.
- Working knowledge of LLM inference: serving engines (e.g. vLLM, SGLang, Triton, Dynamo, production-stack), GPUs, and the performance levers that matter (latency, throughput, memory/KV cache).
- Ability to design a clean benchmark and defend the numbers: same methodology, reproducible results.
- Excellent customer-facing communication: you can explain a tradeoff to an engineer and calm a stakeholder during downtime with equal ease.
- Self-directed and comfortable owning ambiguous problems solo, while documenting so the next person doesn't start from scratch.
Bonus points
- Experience with KV-cache optimization, caching systems, or distributed inference.
- On-prem / air-gapped / regulated-environment deployment experience.
- Familiarity with LMCache or similar open-source inference-optimization frameworks.
- Experience building internal tooling or playbooks that scaled a services/deployment team.
Location & logistics
- US-based, hybrid in Foster City, CA or remote.
Pay: $100,000.00 - $200,000.00 per year
Benefits:
- 401(k)
- 401(k) matching
- Dental insurance
- Dependent health insurance coverage
- Health insurance
- Health savings account
- Life insurance
- Paid holidays
- Unlimited paid time off
- Visa sponsorship
- Vision insurance
Work Location: Remote
About Forward Deployed Engineering
Forward Deployed Engineers are embedded directly with customers to build custom solutions, integrate products into existing infrastructure, and bridge the gap between product engineering and customer success. The role combines deep technical skills with the ability to operate in client environments and translate business requirements into working software.
Originally pioneered by Palantir, the FDE model has spread across AI, enterprise SaaS, and cloud infrastructure companies. FDEs write production code, architect integrations, train customer teams, and feed product insights back to the core engineering organization. At companies like OpenAI, Salesforce, and Databricks, FDE teams are treated as elite engineering units that can ship custom solutions in days rather than quarters.
Typical FDE stack: Python, TypeScript, SQL, REST/GraphQL APIs, cloud platforms (AWS/GCP/Azure), and increasingly LLM APIs and AI orchestration frameworks. Strong communication and the ability to context-switch between technical and business conversations are as important as coding ability.
Similar Roles
Get the FDE Pulse Brief
Weekly market intelligence for Forward Deployed Engineers. Job trends, salary data, and who's hiring. Free.