Improved A/B testing acceleration methods for parametric hypothesis testing: T-test comparison with CUPED, CUPED++ and Bayesian Estimator
Abstract
The study aimed to compare statistical analysis methods to improve the testing of alternatives. The study evaluated four main methods: the classic T-test, the conventional and advanced method of Controlled Experiments Using Pre-Experimental Data (CUPED), and the Bayesian Estimator. The main results included a demonstration of the A/B testing process, and the described statistical analysis methods included detailed characteristics and examples of use. The simulations and practical application revealed that the T-test provides high accuracy with small samples, but its effectiveness decreases with increasing sample size due to high resource requirements. The calculator for this method demonstrated effectiveness in simple tasks but had limitations with large data. The conventional CUPED method has shown increased accuracy due to variation correction, but its effectiveness decreases when working with large and complex data sets. The written program for this method has shown to be effective in cases the previous data is well represented, but its capabilities are limited when processing large data sets. The improved version provided a significant improvement in both accuracy and processing speed, especially for large datasets, thanks to advanced modelling and optimisation. The code results confirmed that this method is highly efficient for complex experiments, particularly when processing large amounts of data. Moreover, the Bayesian Estimator demonstrated high accuracy due to the integration of prior knowledge but required more computational resources and time. The platform used for this method demonstrated the ability to account for uncertainty yet required complex model settings. The results highlighted the importance of ing the appropriate statistical analysis method depending on the scale and complexity of the data to ensure optimal accuracy and efficiency of testing. Метою дослідження було порівняння методів статистичного аналізу для покращення тестування альтернатив. У дослідженні оцінювалися чотири основні методи: класичний Т-тест, традиційний і вдосконалений метод контрольованих експериментів з використанням передекспериментальних даних (CUPED), а також Байєсівський оцінювач. Основні результати включали демонстрацію процесу A/B тестування, а описані методи статистичного аналізу включали детальні характеристики та приклади використання. Моделювання та практичне застосування показали, що T-тест забезпечує високу точність при невеликих вибірках, але його ефективність знижується зі збільшенням розміру вибірки через високі вимоги до ресурсів. Калькулятор для цього методу продемонстрував ефективність у простих завданнях, але мав обмеження при роботі з великими даними. Традиційний метод CUPED показав підвищену точність завдяки варіаційній корекції, але його ефективність знижується при роботі з великими та складними наборами даних. Написана програма для цього методу показала свою ефективність у випадках, коли попередні дані добре представлені, але її можливості обмежені при обробці великих масивів даних. Вдосконалена версія забезпечила значне покращення як точності, так і швидкості обробки, особливо для великих наборів даних, завдяки вдосконаленому моделюванню та оптимізації. Результати роботи коду підтвердили, що цей метод є високоефективним для складних експериментів, особливо при обробці великих обсягів даних. Крім того, Байєсівський оцінювач продемонстрував високу точність завдяки інтеграції попередніх знань, але вимагав більше обчислювальних ресурсів і часу. Платформа, що використовувалася для цього методу, продемонструвала здатність враховувати невизначеність, але вимагала складних налаштувань моделі. Результати підкреслили важливість вибору відповідного методу статистичного аналізу залежно від масштабу та складності даних для забезпечення оптимальної точності та ефективності тестування
URI:
https://ir.lib.vntu.edu.ua//handle/123456789/52382

