Cultivate

Identifying employee relationships from emails for improved HR analytics: Cultivate consulting project, Insight DS.17C

Data ScienceJupyter NotebookMIT

Abstract

Cultivate is an open-source Data Science project. Identifying employee relationships from emails for improved HR analytics: Cultivate consulting project, Insight DS.17C. During my time as an Insight Data Science Fellow, I consulted with Cultivate, a SaaS business whose mission is to provide human resource (HR) departments with a better way to measure and evaluate employees. Prior to joining Insight, I received my Ph.D. It is built using Jupyter Notebook. Key capabilities include: Messages without a recipient; Messages where the sender and recipient were the same; Messages not from an @enron.com address, as these generally represented spam or subscription emails. The complete source code is publicly available on GitHub under the MIT License, making it a useful reference for students building a Data Science mini project or final-year project.

1. Introduction

During my time as an Insight Data Science Fellow, I consulted with Cultivate, a SaaS business whose mission is to provide human resource (HR) departments with a better way to measure and evaluate employees. Prior to joining Insight, I received my Ph.D. from the Program in Neuroscience at Harvard. As I prepared for my transition to industry, I became fascinated by how large organizations manage and motivate employees so that teams can be most productive. I was therefore excited by the opportunity to work with Cultivate to learn about employee communications. Links to my slides can be found here and my webapp can be found here.

Due to privacy concerns, Cultivate could not share their client companies’ data; however, a widely used dataset to mine insights about email communication is the Enron email database. Following the 2001 Enron scandal and federal investigations, many emails of Enron employees were posted on the web, and this dataset is now publicly available from Carnegie Mellon’s Computer Science department. I downloaded the dataset and converted it from mbox to json format. Although there are ~500,000 total messages in the Enron dataset, many of these seemed to be draft emails or other miscellaneous digital documents. To extract insight about employee relationships, I downloaded the emails in the _sent_mail, discussion_threads, inbox and sent_items folders, as these were common across all email senders in the database and likely encompass the majority of meaningful communications. In total, there were >12,000 emails in these folders.

2. Objective

Identifying employee relationships from emails for improved HR analytics: Cultivate consulting project, Insight DS.17C

This project demonstrates how Jupyter Notebook can be applied to a real-world Data Science problem.

3. Key Features / Modules

  • Messages without a recipient
  • Messages where the sender and recipient were the same
  • Messages not from an @enron.com address, as these generally represented spam or subscription emails
  • Messages from identifiable administrative accounts
  • Duplicate messages
  • Empty messages
  • Messages where Fw or Fwd appeared in the Subject field
  • Messages from infrequent message senders (i.e., senders who only sent 1-2 emails in the database)

4. Technology Stack

Jupyter Notebook

5. System Requirements

General requirements for this technology stack — check the README for exact versions.

  • Python 3.8 or later with Jupyter Notebook / JupyterLab (or Google Colab)
  • pip for dependencies
  • Git (to clone the repository)

6. Installation & Setup

git clone https://github.com/yuwie10/cultivate.git
cd cultivate

Full setup instructions are in the project README.

7. Future Enhancements

Suggested extensions you can add to make this your own project.

  • Turn the analysis into an interactive dashboard
  • Automate data refresh with a scheduled job
  • Add a predictive model on top of the analysis

8. Viva / Review Questions

Common questions examiners ask for projects in this domain.

  1. What is the source of the dataset and how was missing data handled?
  2. Which exploratory analysis steps revealed the most useful insight?
  3. Why were these particular charts chosen to present the data?
  4. Which statistical or ML technique supports the conclusions?
  5. How could the analysis be automated or refreshed with new data?

9. Source Code & License

This project is developed by yuwie10 and published on GitHub under the MIT License. Please follow the license terms and credit the original author when you use or modify this code.

Want to build this as your internship project?

Work on a Data Science project like this with mentor guidance, weekly reviews and an internship certificate from Training Trains, Erode — online or offline.

Apply for Data Science Internship