Consumer brand
80.8%
Recall on churn prediction
Flags customers a median of 4.2 days before they churn. A share of sessions is never scored, so lift is measured against a holdout.
Research
Predefined behavior for general testing. Then proprietary models trained on your real product data, so synthetic customers align with real ones, and every claim is checked on sessions the model never saw.
Six stages, the first of which works without any of your data. Each later stage replaces an assumption with something measured.
01
Behavioral patterns already learned from earlier testing, adapted with your brand context and customer metadata.
Useful for testing general hypotheses. Not yet a representation of your actual customers.
02
Website and app visits, through the SDK, a tag or your warehouse. Anonymous telemetry is enough to model behavior.
03
Segments are found automatically. Labels like “the gift buyer” or “the returning considerer” are names given to segments that emerge from the data, not a fixed list. Other profiles are not only possible but expected.
04
One model per segment of how it navigates every section of the product. Training is automatic and usually takes a couple of days. The same models serve per-request prediction at runtime.
05
The models pull the prior cohorts toward your customers until synthetic users behave like real ones. After each run, any bucket that drifts from its observed range is corrected, and the correction persists.
06
Historical sessions are held out of training and scored as unseen data, so the fit to reality is measurable. Then a simulated outcome is confirmed or rejected with an A/B split inside the platform, within days.
None of them alone would be enough. Blended, and corrected by the calibration loop, they produce behavior that lands in the range your real cohort produces.
Layer 1
Per-bucket transition probabilities exported from your model. A high-intent agent picks its next action from the distribution real high-intent customers produce.
Layer 2
Each agent carries a persona built from the signals your model weights most, and reasons about each step in context: device, language, session depth, prior behavior.
Layer 3
Where the model says a feature matters, agents attend to it. This is what lets a simulation respond when a change touches something customers care about.
The feature layer
260
behavioral features, first built for clinical movement research
Computed over interaction with software instead of limb motion. Better features are why a model needs a few hundred outcomes, where others need thousands of events per user.
Where agents can run
Agents can be given tools: a browser they drive, search, a price comparator, a cart. Every step, hesitation and dead end is recorded, with each agent's own account of why.
Production models from client deployments. Metrics are from held-out data, not benchmarks.
Consumer brand
80.8%
Recall on churn prediction
Flags customers a median of 4.2 days before they churn. A share of sessions is never scored, so lift is measured against a holdout.
B2C ed-tech
93.6%
Accuracy on onboarding abandonment
49,718 held-out sessions. 89.2% of abandoners were flagged before the steps that lose people.
Fashion e-commerce
97.3%
Precision on cart to checkout
2,630 held-out sessions scored at add-to-cart, 83.0% accuracy. Only 1.2% of sessions received an offer they did not need.
Production model
0.942
AUC-ROC on a live client holdout
91.3% accuracy, 0.834 precision, 0.867 recall, 0.96 calibration, measured on held-out sessions rather than a benchmark.
Content simulation · 8 social creatives · 12 runs
Simulation repeatability
9–10%
Tap likelihood for the same creative across four independent runs, while raw taps ranged from 1 to 4.
With a few dozen agents, raw tap counts move with luck. Tap likelihood does not, which is why reports compare creatives on it.
Before product behavior, the same feature construction read movement in clinical studies. Further publications on behavioral simulation are in preparation.
paper · 2024
0.883 AUC · 82.1% agreement with electroneurography
Where the feature layer comes from. Movement captured during six standard neurological tests was turned into behavioral features and scored by an ML model for diabetic neuropathy, then compared with electrodiagnostic examination, the clinical reference, in a 23-participant pilot.
Biosensors 14(4):166 · MDPI · Q1 · indexed in PubMed
study
The same movement features, used by clinicians without an engineering team to screen balance in fibromyalgia patients, with the aim of fewer falls and fractures.
With the Institute of Rheumatology · presented to EULAR
As a contractor, the team also built algorithms for wearable research on detecting infection before symptoms. Read about the study.
Notes on behavior and the brain
Backed by a team of ML and software engineers, building and shipping the platform day to day.

Co-founder & CEO
Led a team delivering products used by hundreds of millions of people.
LinkedIn
Walk your team through the methodology.
Request a briefing