- Sector
- AI
- City
- Bengaluru
- Area
- Koramangala and Outer Ring Road
- Experience
- 1 to 3 years
- Role family
- Data and ML
- Salary
- Not disclosed
- Posted
- 18 Aug 2026 · 2 weeks ago
- Last checked
- 7 Sept 2026
About the role
Emergent builds autonomous coding agents that replace traditional software development by generating, testing, and deploying production applications directly from plain-language intent. Our systems run in production at global scale and are used to build millions of real applications.
Since our public launch, we've crossed $130M in Annualised Revenue and grown to over 10M users across 190+ countries, who have built 12M+ applications on Emergent. We're backed by Creaegis, Khosla Ventures, SoftBank, Lightspeed, Together, Y Combinator, Google, Claypond and Sentinel Global.
We're solving the hard part of AI-driven software creation: correctness, reliability, security, and scale in real production systems. The team is built by repeat founders, Olympiad medalists, IIT & IIM alumni, and leaders from Google, Amazon, and Dropbox.
We're hiring builders who want ownership, speed, and impact at global scale.
The Role:
Every part of our growth runs on decisions made from data: who activates, who converts, what an agent-built app actually costs us to produce, and which users are here to build versus here to abuse the free tier. What makes analytics here unusual is the shape of that data. Alongside the standard event and revenue tables, we sit on an enormous volume of unstructured signal, including millions of agent trajectories (the step-by-step reasoning, actions, and observations of every build), support tickets, HITL feedback, and the natural-language prompts users write. The richest insights in the business are buried in that text, and the person in this role is the one who gets them out. You own the loop end-to-end: what we measure, how we prove it, what the number supports, and what it doesn't.
What You'll Do
- Turn agent trajectories, support tickets, logs, and user prompts into structured, queryable signal through summarize-then-embed-then-cluster pipelines (à la Anthropic's Clio and Braintrust Topics): distill each trace along a dimension with an LLM, embed the summary, cluster and name the patterns, then classify at scale
- Surface early indicators of confusion, a coming bug wave, churn risk, or fraud that no dashboard would ever surface on its own, and route them to the right team
- Build predictive models that forecast conversion, retention, expansion, and churn, and embed those signals directly into product and growth workflows
- Own marketing attribution and MMM: build the media-mix and incrementality models that tell us what's actually driving signups and paid conversions when per-user attribution is partial and, on mobile, broken by design
- Own product and growth analytics across the self-serve funnel, web and mobile, covering activation, engagement, retention, and conversion, and design and analyze A/B and growth tests with real rigor around power, novelty effects, interference, and causal inference
- Run clustering pipelines over hundreds of thousands of agent trajectories to discover the recurring kinds of things users try to build and the recurring ways builds fail, then hand product a taxonomy nobody had to hand-label, along with which clusters predict churn
- Model the "aha moment" for new users, including text-derived features from their first prompts and first agent interactions, and rebuild onboarding around the earliest signals of long-term retention
- Build a gross-margin model that attributes LLM and compute cost down to the individual app and cohort, and tell product which segments are net-positive
- Untangle a fraud ring that looks anomalous on compute spend but has real payment history, decide whether it's an enforcement problem or a pricing problem, and defend the call with the data
The rest of this description is on the employer’s own page.