
Real job — pulled straight from Zenskar’s careers page · Verified August 14, 2026 · No reposts.
Job description
Zenskar is hiring a Staff Engineer, Generative AI — a full-time, based in Bengaluru, India role. Apply directly on Zenskar's careers page below.
Staff Engineer (GenAI)
Location: Bengaluru, India
Department: Engineering
Experience: 8+
About Zenskar
Funding
The Problem We're Solving
What Customers Say
"We're saving 200+ hours/quarter on invoicing and receivables by completely automating our recurring billing."
— Noy Kalansky, Finance Controller, Pontera
"Zenskar automates revenue recognition accurately for our value-based billing: agents reducing manual hours by 70%."
— Matt Barnard, VP Finance, Vertice
"Zenskar’s agents automated 90% of our billing, integrated with our CRM, and accelerated revenue collection by a month."
— Ming Lui, VP Finance, Yembo
"Sardine had spent 4 years running billing in-house for high-volume, usage-based pricing. Zenskar took care of it all."
— Sardine team
"We launched our product 4 months faster instead of building an in-house system for our usage-based pricing."
— Kshitij Gupta, CEO, 100ms
About the role
What you'll do
- Turn ambiguous product asks into working LLM-based features: prompt and context design, tool calling, orchestration across model calls, structured output handling
- Design for the failure modes non-deterministic systems actually have: hallucination, drift, silent wrongness. Build guardrails, and know the difference between demo-grade and production-grade because you've shipped the gap between them
- Make sure an LLM is never the unverified source of truth for a financial decision. Design the verification or human-in-loop layer wherever money is actually at stake
- Pick the right model or vendor for the job (cost, quality, latency, data policy) and avoid lock-in. Once chosen, run it efficiently: token economics, caching, batching, per-request cost tracking
- Set the technical standard other teams build their Gen AI features against, and get multiple teams to actually adopt it without having direct authority over them
- Build and own the eval harness that catches quality regressions before they ship, not ad hoc spot-checks
- Know when not to use an LLM at all, where a deterministic system would be cheaper, more reliable, or simply correct
- Name and quantify technical debt in Gen AI systems (prompt sprawl, eval gaps, untracked cost, brittle integrations) and negotiate time to address it
- Shape who joins the team by holding a technical bar in interview loops, without owning headcount decisions
What drives you
- You've seen AI systems fail in production and got obsessed with understanding why
- You think about reliability, evaluation, and observability before you think about model selection
- You believe probabilistic systems deserve the same engineering rigor as deterministic ones
- You care about systems users can trust, not systems that look impressive in a demo
- Gen AI's baseline moves monthly for you, not yearly. You actively track it: specific people, sources, releases, and can point to a real decision that changed because of a recent shift
What you'll need
- 8+ years of software engineering experience, including meaningful experience shipping AI-powered products to production
- A real, shipped-at-scale system in your track record, doesn't have to be Gen AI. If it's ML work, it shipped to real users, not fine-tuning or research that stayed in a notebook
- Deep understanding of LLM application architecture: tool use, structured outputs, retrieval, orchestration, and where these actually break in production
- Experience building agentic systems that do multi-step reasoning and interact with tools or business systems
- A track record of treating prompts as versioned, testable, deployable artifacts, not strings scattered through the codebase
- Strong RAG fundamentals: retrieval quality, chunking, embeddings, evaluation, knowledge system design
- Experience designing evaluation frameworks and regression pipelines for AI systems
- Strong grasp of AI observability: tracing, monitoring, feedback loops, production debugging
- A track record of getting more than one team to adopt a pattern or standard you set, without owning those teams
- Judgment for when not to use an LLM at all
- Strong backend engineering skills, and enough frontend ability to own an AI experience end to end
- Can clearly walk through an AI system you've built two levels below the pitch: what failed, what you learned, how you improved reliability over time
Good to have
- Memory architectures: episodic memory, procedural memory, retrieval systems, knowledge stores
- Agent orchestration frameworks (LangGraph, Pydantic AI, OpenAI Agents SDK, CrewAI, or a custom runtime)
- Long-running autonomous workflows and event-driven agent systems
- Voice, multimodal, or real-time AI systems
- Fine-tuning experience and understanding of model internals beyond API consumption
- Open-source model deployment and inference infrastructure
- Experience with financial systems, billing platforms, revenue operations, accounting, or fintech
- Familiarity with MCP, tool ecosystems, and AI platform architecture
- Startup experience and comfort operating with high ownership and no formal authority
Location
- Hybrid - 2 days per week
- Office Location: Indiranagar, Bengaluru.
- Address: 3rd Floor, A wing No 1, Carlton Towers, HAL Old Airport Rd, HAL 2nd Stage, Indiranagar, Bengaluru, Karnataka 560008.
Interview process
- R0, Recruiter screen (30 min): Fit, motivation, and a quick check that you have real, specific experience behind your background, not generic answers.
- R1, Technical: critique a broken Gen AI integration (60 min): You'll be handed a small billing-shaped Gen AI integration with seeded defects and asked to find them, including the one that would go unnoticed until it was already wrong on a customer's invoice.
- R2, Technical: design a new integration and its eval framework (60 min): An intentionally underspecified product ask. You design the integration, decide whether an LLM belongs in it at all, choose a model with real reasoning, and design how you'd know it's working before a customer complains.
- R3, Project audit (60 min): A numbers-anchored walkthrough of one real system you've shipped at real scale, architecture, usage, cost, and the debt it accumulated.
- R4, Architect and influence (60 min): Whether you've actually set a technical pattern that more than one team adopted, and how you've held a hiring bar without owning headcount.
- R5, Bar raiser (60 min): Values, judgment, and a deeper probe on whatever came through thinnest earlier in the process.
- Reference checks: Two former colleagues, run by two different people on our side, including a deep dive on a Gen AI system you shipped and what happened after it hit production.
How to apply
Get Staff Engineer, Generative AI jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs

Distinguished Software Engineer, Data Platform



Frequently asked questions
What skills are required for Staff Engineer, Generative AI at Zenskar?
The required skills for Staff Engineer, Generative AI at Zenskar include: Generative AI, LLM, Python, LangGraph.
What is the seniority level for Staff Engineer, Generative AI at Zenskar?
Staff Engineer, Generative AI at Zenskar is a Staff / Senior level position.
How do I apply for Staff Engineer, Generative AI at Zenskar?
You can view the full description and apply for Staff Engineer, Generative AI at Zenskar on EchoJobs: https://echojobs.io/job/zenskar-staff-engineer-genai-0evmp.