Pyprocessmacro

A Python library for moderation, mediation and conditional process analysis.

Data SciencePythonMIT

Abstract

Pyprocessmacro is an open-source Data Science project. A Python library for moderation, mediation and conditional process analysis. PyProcessMacro was developed by Quentin André during his PhD in Marketing at INSEAD Business School, France. It is built using Python. Key capabilities include: All models (1 to 76), with the exception of Model 6 (serial mediation) are supported, and have been numerically; Estimation of binary/continuous outcome variables. The binary outcomes are estimated in Logit using the; All statistics reported by Process:. The complete source code is publicly available on GitHub under the MIT License, making it a useful reference for students building a Data Science mini project or final-year project.

1. Introduction

PyProcessMacro was developed by Quentin André during his PhD in Marketing at INSEAD Business School, France.

PyProcessMacro: A Python Implementation of Andrew F. Hayes' 'Process' Macro

Because PyProcessMacro is a complete reimplementation of the Process Macro, and was not based on the original code, permission was generously granted by Andrew F. Hayes to distribute PyProcessMacro under a MIT license.

2. Objective

A Python library for moderation, mediation and conditional process analysis.

This project demonstrates how Python can be applied to a real-world Data Science problem.

3. Key Features / Modules

  • All models (1 to 76), with the exception of Model 6 (serial mediation) are supported, and have been numerically
  • Estimation of binary/continuous outcome variables. The binary outcomes are estimated in Logit using the
  • All statistics reported by Process:
  • Variable parameters for outcome models
  • (Conditional) direct and indirect effects
  • Indices for Partial/Conditional/Moderated Moderated Mediation are always reported if the model supports them.
  • Automatic generation of spotlight values for continuous/discrete moderators.
  • Rich set of options to tweak the estimation and display of the different models: (almost) all the options from
  • Variable names can be of any length, and even include spaces and special characters.
  • All mediation models support an infinite number of mediators (versus a maximum of 10 in Process).

4. Technology Stack

Python

5. System Requirements

General requirements for this technology stack — check the README for exact versions.

  • Python 3.8 or later
  • pip / virtualenv for dependencies
  • VS Code, PyCharm or Jupyter Notebook
  • Git (to clone the repository)

6. Installation & Setup

git clone https://github.com/QuentinAndre/pyprocessmacro.git
cd pyprocessmacro
  1. On the x-axis (moderator passed to x).
  2. As a color-code, in which case several lines are displayed on the same plot (moderator passed to hue).
  3. On different plots, displayed side-by-side (moderator passed to col).
  4. On different plots, displayed one below the other (moderator passed to row)
from pyprocessmacro import Process
import pandas as pd
import matplotlib.pyplot as plt

df = pd.read_csv("MyDataset.csv")
p = Process(data=df, model=13, x="Effort", y="Success", w="Motivation", z="SkillRelevance",
            m=["MediationSkills", "ModerationSkills"], suppr_init=True)

# Conditional direct effects of Effort, at values of Motivation (x-axis)
g = p.plot_direct_effects(x="Motivation")
plt.show()
# Conditional indirect effects through MediationSkills, at values of Motivation (x-axis) and
# SkillRelevance (color-coded)
g = p.plot_indirect_effects(med_name="MediationSkills", x="Motivation", hue="SkillRelevance")
g.add_legend(title="") # Add the legend for the color-coding
plt.show()
# Display the values for SkillRelevance on side-by-side plots instead.
g = p.plot_indirect_effects(med_name="MediationSkills", x="Motivation", col="SkillRelevance")
plt.show()
# Display the values for SkillRelevance on vertical plots instead.
g = p.plot_indirect_effects(med_name="MediationSkills", x="Motivation", row="SkillRelevance")
plt.show()

Full setup instructions are in the project README.

7. Future Enhancements

Suggested extensions you can add to make this your own project.

  • Turn the analysis into an interactive dashboard
  • Automate data refresh with a scheduled job
  • Add a predictive model on top of the analysis

8. Viva / Review Questions

Common questions examiners ask for projects in this domain.

  1. What is the source of the dataset and how was missing data handled?
  2. Which exploratory analysis steps revealed the most useful insight?
  3. Why were these particular charts chosen to present the data?
  4. Which statistical or ML technique supports the conclusions?
  5. How could the analysis be automated or refreshed with new data?

9. Source Code & License

This project is developed by QuentinAndre and published on GitHub under the MIT License. Please follow the license terms and credit the original author when you use or modify this code.

Want to build this as your internship project?

Work on a Data Science project like this with mentor guidance, weekly reviews and an internship certificate from Training Trains, Erode — online or offline.

Apply for Data Science Internship