A/B Test Report

Overview

The A/B Test Report provides cohort-based analytics for A/B tests configured in Magify. It helps you evaluate how changes in your application — such as design, algorithms, or pricing — affect user behavior, retention, and monetization.

The report can also provide daily analytics. However, the primary date used to select users is the date when they were assigned to the test. This means that even when analyzing data by day, the report includes only users who entered the selected A/B test during the specified assignment dates.

The A/B Test Report includes statistics only for users who were assigned to the selected test.

How the Report Works

The A/B Test Report is cohort-based. A cohort is formed based on the date when a user was assigned to an A/B test and test group (Assign Date).

Activity, revenue, and retention metrics are calculated for these users after assignment. The event date (Event Date) is secondary to the cohort definition, but remains available for filtering and charting.

This allows you to answer the main question of an A/B test: Which group performs better after the tested changes are applied?

Assign Date

Assign Date is the date when a user was included in an A/B test and assigned to one of its groups. It defines the cohort used by the report.

The report uses different dates as the basis for different types of analytics:

  • Daily Report answers: "What happened on January 1?" Its basis is Event Date.
  • UA Report answers: "How did users who installed the application on January 1 behave on Day 1, Day 2, or Day 7 after installation?" Its basis is Install Date.
  • A/B Test Report answers: "How did users assigned to a test on January 1 behave on Day 1, Day 2, or Day 7 after assignment?" Its basis is Assign Date.

For example, if a user installed the game on December 1 but entered the test on December 15, December 15 is Day 0 for this report. The user's activity from December 1 through December 14 is not included in the cohort that started on December 15.

This is intentional: the report is designed to evaluate the impact of the test rather than the user's previous experience.

Assignment logic

Test assignment data is collected by the Magify SDK.

For campaign-based tests, the assignment date is determined by the actual impression of the test campaign after the user receives the test configuration. In this case, the user's experience changes only after the test is actually shown.

For feature-based tests, the assignment date is the moment when the user receives the test configuration, without requiring an additional event. Receiving the test features is considered sufficient to start tracking the user's subsequent behavior.

Assign Date and Event Date

Event Date is the date when an event occurred.

Days After Assign is the number of days between the user's assignment to the test and the event, such as an ad impression, purchase, or application launch.

Use Days After Assign as the X-axis when you need to build cohort charts showing how users behave after assignment.

For example, select:

  • Assign Date = 2024-01-01
  • Days After Assign = 0,1,2,3,4,5,6,7
  • Metric = Revenue per User

The resulting chart shows how the average revenue per user from the January 1 cohort changed during their first week in the test.

Retention

Retention in the A/B Test Report is calculated from Assign Date, not from the application installation date.

This is important because A/B tests can include both new and existing users. For an existing user, Retention D1 means that the user returned to the application one day after the test was applied to them.

Retention is calculated using the cohort approach:

  • The denominator includes all users assigned to the test on Day X.
  • The numerator includes users from that cohort who were active on Day X+N.

Data Sources

The A/B Test Report combines data from several sources:

  • Test assignment data — collected by the Magify SDK. The assignment logic depends on whether the test is campaign-based or feature-based, as described above.
  • Activity data — collected by the Magify SDK from user devices.
  • Purchases and refunds — received from the App Store and Google Play. Purchase validation must be configured.
  • Install data — received from the MMP grabber and used to enrich user data.

Purchase and subscription data is collected and/or validated through the application stores, while activity and advertising data comes from the Magify SDK.

All monetary values are converted to USD for comparison. Store commission is deducted from purchases.

Getting Started

To start analyzing an A/B test:

  1. Select a specific A/B test using the Assign A/B Test filter.
  2. Add Assign A/B Group to Group by to view the test groups side by side.
  3. Build a chart using DAU and/or Total Revenue per User as the metric and Days After Assign as the X-axis (Order by).
  4. Compare the curves for the Control and Test groups to see which group has higher activity or revenue.
  5. For tests with a specific goal, select the metric that corresponds to the test objective and research plan.

Assign A/B Test is a required filter in the report. We strongly recommend viewing statistics with Assign A/B Group included in the breakdown to compare test groups correctly.

Use Days After Assign as the X-axis to track how the cohort behaves after assignment.

!!!!!!!!!!! АААААА

Dimensions and Filters

Dimensions define how report data can be filtered, grouped, and broken down for analysis.

Metrics

Metrics provide quantitative data for evaluating A/B test performance, user activity, retention, monetization, and other user behavior.

What You Can Analyze with This Report

The A/B Test Report can be used to compare test groups across user activity, retention, monetization, gameplay behavior, and traffic sources.

Getting Started

  1. Select a specific A/B test using the Assign A/B Test filter.
  2. Add Assign A/B Group to Group by to compare data for different test groups side by side.
  3. Create a chart using DAU and/or Total Revenue per User as the metric and Days After Assign as Order by on the X-axis.
  4. Compare the curves for the Control and Test groups to see which group shows higher activity or revenue. This provides a basic view of the test's performance.
  5. For tests with a specific goal, select the metric that corresponds to the test objective and research plan.

A/B Test Performance Analysis

Use this analysis when deciding whether a tested change should be rolled out to all users.

Key metrics:

  • Assigned Users — check that the test groups are comparable in size.
  • Retention D1, Retention D3, Retention D7, Retention D30 — evaluate the impact of the test on user retention.
  • Revenue per User — measure the impact on monetization.
  • Sessions per User, Playtime per Session — evaluate the impact on engagement.

Recommended dimensions:

  • Days After Assign — X-axis.
  • Assign A/B Group — lines on the chart.

Monetization Analysis

Use this analysis to understand how a tested change affects purchasing behavior.

Key metrics:

  • IAP Revenue, Ad Revenue
  • IAP Purchases Completed, Paid Sub Activations
  • Average IAP Purchase Value

Recommended dimensions:

  • Days After Assign
  • Product, Product Type — identify which products perform better or worse between test groups.
  • Country — identify countries where the test performs better.

Gameplay Analysis

Use this analysis to understand how a tested change affects player behavior.

Key metrics:

  • Levels Started, Levels Completed, Levels Failed
  • Level Complete-to-Fail Ratio
  • Boosters Spent, Boosters Received

Recommended dimensions:

  • Level
  • Game Mode
  • Bonus Type
  • Bonus Source

For example, if a test changes the difficulty of a specific level, use Level to compare completion and failure metrics between the Control and Test groups.

Traffic Source Analysis

Use this analysis to understand whether the test performs differently for users acquired from different traffic sources.

Key metrics:

  • Retention D1, Retention D7
  • Revenue per User
  • IAP Revenue
  • Ad Revenue

Recommended dimensions:

  • Install Media Source
  • Install Campaign Name
  • Install Ad Set Name
  • Install Ad Name
  • Country

This analysis helps identify whether the test effect differs depending on the user's acquisition source.

Technical Considerations and Important Notes

Cohort Basis: Assign Date

The A/B Test Report is based on Assign Date — the date when a user was assigned to an A/B test and test group.

This is the key difference between the A/B Test Report and other report types:

  • Daily Report answers: "What happened on January 1?" Its primary date is Event Date.
  • UA Report answers: "How did users who installed the app on January 1 behave on Day 1, Day 2, Day 7, and so on after installation?" Its cohort is based on Install Date.
  • A/B Test Report answers: "How did users assigned to the test on January 1 behave on Day 1, Day 2, Day 7, and so on after assignment?" Its cohort is based on Assign Date.

For example, if a user installed the game on December 1 but was assigned to the test on December 15, December 15 is Day 0 for that user in the A/B Test Report. Activity between December 1 and December 14 is not included in the cohort that starts on December 15.

This is intentional: the report is designed to evaluate user behavior after the test is applied rather than the user's previous experience.

Retention Metrics

Retention in the A/B Test Report is calculated from the date when the user was assigned to the test (Assign Date).

This is important because A/B tests can target both new and existing users. For an existing user, Retention D1 means that the user returned to the application one day after being assigned to the test.

Retention follows a cohort-based calculation:

  • Denominator — all users assigned to the test on Day X.
  • Numerator — users from that cohort who were active on Day X + N.

Assigned Users vs Assign Events

Assigned Users and Assign Events measure different aspects of assignment to an A/B test:

  • Assigned Users counts unique users assigned to the test.
  • Assign Events counts assignment events.

An assignment event (assign_events) is generated when a user meets the test conditions. If user identification is disrupted — for example, if the client_id is reset — the user may not be included in Assigned Users because the report cannot identify them correctly.

A significant difference between Assigned Users and Assign Events may therefore indicate an issue with user identification on the device.

Days After Assign

Use Days After Assign on the X-axis to build cohort-based charts showing how users behave after being assigned to the test.

For example, select:

  • Assign Date = 2024-01-01
  • Days After Assign = 0, 1, 2, 3, 4, 5, 6, 7
  • Revenue per User as the metric

The resulting chart shows how the average revenue per user for the January 1 cohort changes during the first week after assignment to the test.

Refunds and Renewals

The same rules used in the Daily Report apply to refunds and renewals in the A/B Test Report:

  • In overall totals for IAP Revenue, Subscription Revenue, and Total Revenue, refunds are deducted.
  • When data is broken down by dimensions other than Application, Date, Country, Assign A/B Test, Assign A/B Group, and Product, refunds are not deducted because app stores do not provide this metadata.

For store-based metrics, whether an event is included in the report also depends on its Event Date. The event date must be after the date when the user was assigned to the test.

Data Freshness

The report includes data when the difference between the event time and the time the event is received in the database does not exceed 7 days.

This reconciliation window keeps A/B test results up to date, but events received with a delay of more than 7 days — for example, events from users who remained offline — are not included in the report.

A/B Test Statistics Guide

This section summarizes the key statistical concepts required to correctly plan, run, and analyze A/B tests. It is intended to help you avoid common mistakes and make informed decisions based on the data available in this report.

Key idea: An A/B test is not simply a comparison of two numbers. It is a scientific experiment governed by statistical principles. Without understanding these principles, random fluctuations can be mistaken for meaningful results and lead to incorrect product decisions.

Key Statistical Concepts

How This Relates to the Report

The report provides the data but does not make statistical conclusions for you. Apply the statistical principles described in this section when interpreting the report results.

Common Mistakes

1. Trusting the Result at Face Value

Mistake: Group B is 10% better, so you assume the result is real.

Reality: The numbers alone do not prove that the difference is not random. Use P-Value to evaluate statistical significance.

How to use the report: Always check the sample size using Assigned Users. With a small sample, even a 50% improvement may be noise.

2. Treating Test Results as Permanent

Mistake: The test "proves" that Group B will always perform better.

Reality: The result shows that an improvement occurred during the test for this group of users.

How to use the report: Make sure the users included in the test have an acquisition structure similar to the overall audience, including country and traffic source.

While the test is running, avoid major changes in user acquisition, promotions, or external advertising campaigns. If you are planning a product change other than the one being tested, run the test after that change.

If the test was conducted under conditions that could have affected the result, repeat it after a month. If the result is reproduced, confidence in it is higher.

3. Overinterpreting Statistical Significance

Mistake: Users "preferred" Group B.

Reality: The test measures how the change affected user behavior, not what users think about it.

How to use the report: Do not infer reasons from the report data. The report shows what happened, but not why. Additional research is required to understand the reasons.

How to Handle a Non-Significant Result

If the calculator shows that the result is not statistically significant (P-Value > 0.05), there are three possible approaches.

1. Get Consistent Data

Problem: External factors affected the results, such as holidays or the release of a new TV series season.

Solution: Exclude these periods from the analysis or run the test again during a more stable period.

How to use the report: Use dimensions to filter out periods affected by external factors.

2. Aim for a Larger Effect

Problem: You are trying to detect an effect that is too small, for example, a +0.5% increase in conversion.

Solution: Make the tested change more substantial. Instead of testing a button color, test a new purchase flow.

How to use the report: Before starting the test, use a calculator to estimate the sample size required to detect the expected effect. If you cannot collect enough users within a reasonable period, the tested change may be too small.

3. Increase the Test Duration and Sample Size

Problem: There is not enough data or the metric has high variance.

Solution: Run the test for longer. A weekly cycle is a good minimum.

How to use the report: Monitor Assigned Users. Do not end the test early as soon as you see a favorable result.

Handling Data Variability

Mobile application data is never constant. It can be affected by the day of the week, seasonality, holidays, and news.

To account for this variability:

  1. Run the test for at least 1–2 weeks to cover a complete weekly cycle, including weekends.
  2. Use Days After Assign or Day Of Week as a dimension instead of calendar dates. This helps smooth out uneven traffic across different days.
  3. Avoid starting tests during major holidays or advertising campaigns unless their impact is what you are testing.

A/B Testing Rules

These rules should be followed consistently. The report is the primary tool for applying them.

Rule 1: Revisit Previous Tests

Principle: Review the results of previous tests again after several months.

How to use the report: After 3–6 months, create a new report using the same filters (Assign A/B Test, Assign A/B Group) but for later dates. Compare the new results with the previous ones and check whether the outcome has changed.

Rule 2: Run an A/A Test

Principle: Before testing changes, make sure the system does not detect differences between two identical groups.

How to use the report: Create a test where Group A and Group B receive the same configuration. Assign both groups in the report and filter by Assign A/B Group. Select the target metric, for example, Revenue per User.

If the report shows a statistically significant difference when checked with an external calculator, the test configuration is not working correctly. Fix the configuration before running an actual A/B test.

Rule 3: Follow an Overall Testing Strategy

Principle: Product development should follow a unified strategy, with A/B tests used as a tool for adjustment.

How to use the report: Select one primary target metric, for example, Revenue per User, and evaluate all tests against it.

If a test increases Revenue per User but decreases Retention D30, make the decision based on your product strategy, for example, whether the goal is to focus on high-spending users or a broader audience.

Rule 4: Do Not Peek

Principle: Do not check the results every day or stop the test as soon as P-Value becomes < 0.05. Wait until the required sample size has been reached.

How to use the report: Calculate the required sample size before starting the test. Start the test and do not check the report until Assigned Users reaches that value. Only then export the data and calculate P-Value.

Other Rules

  • Run tests consistently. Do not change test parameters or the target metric while the test is running.
  • Never stop a test early to change its conditions. Doing so invalidates the collected data.
  • Run a test for at least 7 days to account for weekly seasonality.

Practical Workflow: From Hypothesis to Conclusion

  1. Formulate a hypothesis.
    Example: "Changing the Starter Pack price from $5 to $10 will not affect purchase conversion" (H₀). H₁: "It will affect purchase conversion."
  2. Select the target metric.
    For example, IAP Revenue per User 7 days after assignment (Days After Assign = 7).
  3. Calculate the required sample size.
    Use a calculator. You need to know:
    • the current metric value;
    • the expected minimum effect;
    • statistical power (80%);
    • significance level (95%).
  4. Configure the A/B test in Magify.
    Create configurations for Groups A and B.
  5. Start the test and DO NOT PEEK.
  6. Wait until the required number of Assigned Users is reached.
    Use the report to monitor the sample size. Once the required number is reached, export IAP Revenue per User data for each group.
  7. Use a statistical significance calculator.
    Enter the exported data for both groups into the calculator.
  8. Make a decision:
    • P-Value ≤ 0.05 → Reject H₀. The change is statistically significant.
    • P-Value > 0.05 → Do not reject H₀. There is no evidence of an effect. Return to the previous version or revise the hypothesis.

Remember: The report provides the data. Statistics provides the tools to interpret it. Together, they support informed business decisions.

Related articles

User Activity

Ad Revenue

In-App Purchases & Subscriptions

Product Report: Methodology and FAQ

Understanding Parent And Nested Campaigns

Product Report