Skip to content

Research

How Moveo One builds a population that behaves like yours.

Predefined behavior for general testing. Then proprietary models trained on your real product data, so synthetic customers align with real ones, and every claim is checked on sessions the model never saw.

01Method

From anonymous telemetry to a calibrated cohort.

Six stages, the first of which works without any of your data. Each later stage replaces an assumption with something measured.

  1. 01

    Start from prior cohorts

    Behavioral patterns already learned from earlier testing, adapted with your brand context and customer metadata.

    Useful for testing general hypotheses. Not yet a representation of your actual customers.

  2. 02

    Read anonymous telemetry

    Website and app visits, through the SDK, a tag or your warehouse. Anonymous telemetry is enough to model behavior.

  3. 03

    Derive behavioral segments

    Segments are found automatically. Labels like “the gift buyer” or “the returning considerer” are names given to segments that emerge from the data, not a fixed list. Other profiles are not only possible but expected.

  4. 04

    Train probabilistic models

    One model per segment of how it navigates every section of the product. Training is automatic and usually takes a couple of days. The same models serve per-request prediction at runtime.

  5. 05

    Recalibrate the cohorts

    The models pull the prior cohorts toward your customers until synthetic users behave like real ones. After each run, any bucket that drifts from its observed range is corrected, and the correction persists.

  6. 06

    Validate twice

    Historical sessions are held out of training and scored as unseen data, so the fit to reality is measurable. Then a simulated outcome is confirmed or rejected with an A/B split inside the platform, within days.

02Inside an agent

Three signals behind every decision a synthetic customer makes.

None of them alone would be enough. Blended, and corrected by the calibration loop, they produce behavior that lands in the range your real cohort produces.

Layer 1

Learned transitions

Per-bucket transition probabilities exported from your model. A high-intent agent picks its next action from the distribution real high-intent customers produce.

Layer 2

Persona reasoning

Each agent carries a persona built from the signals your model weights most, and reasons about each step in context: device, language, session depth, prior behavior.

Layer 3

Feature importance

Where the model says a feature matters, agents attend to it. This is what lets a simulation respond when a change touches something customers care about.

The feature layer

260

behavioral features, first built for clinical movement research

Computed over interaction with software instead of limb motion. Better features are why a model needs a few hundred outcomes, where others need thousands of events per user.

HesitationReaction speedCognitive loadCorrection rateAttention decayDepth of considerationRhythmWillingness to engage

Where agents can run

  • Your live product, in a real browser or app
  • A rebuild that has not shipped
  • Third-party surfaces: marketplaces, retailers, comparison sites
  • Constructed environments for content: a film, a post, a banner, a pack

Agents can be given tools: a browser they drive, search, a price comparator, a cart. Every step, hesitation and dead end is recorded, with each agent's own account of why.

03Validation

Scored on sessions the model never saw.

Production models from client deployments. Metrics are from held-out data, not benchmarks.

Consumer brand

80.8%

Recall on churn prediction

Flags customers a median of 4.2 days before they churn. A share of sessions is never scored, so lift is measured against a holdout.

B2C ed-tech

93.6%

Accuracy on onboarding abandonment

49,718 held-out sessions. 89.2% of abandoners were flagged before the steps that lose people.

Fashion e-commerce

97.3%

Precision on cart to checkout

2,630 held-out sessions scored at add-to-cart, 83.0% accuracy. Only 1.2% of sessions received an offer they did not need.

Production model

0.942

AUC-ROC on a live client holdout

91.3% accuracy, 0.834 precision, 0.867 recall, 0.96 calibration, measured on held-out sessions rather than a benchmark.

Content simulation · 8 social creatives · 12 runs

CreativeTap likelihood
  • HOffer with conditions listed8–11%
  • AConditional offer, price comparison10%
  • BHeadline offer, question as the hook9–10%
  • COffer with a condition, large numeral9%
  • ESeasonal visual, price drop9%
  • DPerson on camera, question-led6%
  • FStory-led, two people, a handoff5%
  • GSocial-proof counter5%
Offer-first Story-led Individual run

Simulation repeatability

9–10%

Tap likelihood for the same creative across four independent runs, while raw taps ranged from 1 to 4.

With a few dozen agents, raw tap counts move with luck. Tap likelihood does not, which is why reports compare creatives on it.

04Origins

The feature layer was first scored against clinical instruments.

Before product behavior, the same feature construction read movement in clinical studies. Further publications on behavioral simulation are in preparation.

As a contractor, the team also built algorithms for wearable research on detecting infection before symptoms. Read about the study.

05Team

Machine learning and clinical research, then product.

Backed by a team of ML and software engineers, building and shipping the platform day to day.

Nikola Gavranović

Nikola Gavranović

Co-founder & CTO

Led machine-learning research in academia.

LinkedIn

Walk your team through the methodology.

Request a briefing