Skip to main content

Reflecting on My Capstone: Lessons from a Semester of Data Science

This capstone was one of my proudest projects, not because it went perfectly, but because it taught me so much through feedback, mistakes, and revision. Over the semester, I worked through the full process of building a data science project, from choosing a research question to cleaning data, building models, evaluating results, and creating a final report. This post is a behind-the-scenes reflection on what I built, what I struggled with, and what I learned along the way.

In this post:

    - Why This Capstone Mattered
    - Choosing the Topic: From Interest to Research Question
    - Data Cleaning Process
    - Building the Models
    - Learning How to Evaluate Results
    - Turning Assignments into a Final Report
    - Final Reflection: What I Learned

Why This Capstone Mattered

In my final capstone for my data science bachelor’s degree, I researched the connection between pre-existing conditions, simple demographics, and adverse COVID-19 outcomes. For this project, an adverse outcome was defined as ICU admission, being put on a ventilator, or death. I also worked on predicting which positive cases in the dataset would have an adverse outcome based on these predictors.

I had completed smaller projects in other classes and hackathons, where I applied what I learned and presented my results, but this was much bigger than anything I had done before. The capstone included a 15–20 page final report and a 25-minute professional-style presentation to the class and three professors.

The project was spread across the entire semester and broken into five parts: Proposal and Research, Cleaning Data and Exploratory Data Analysis (EDA), Building Models, Evaluating Results, and the Final Report. This class was a valuable experience because it gave me the chance to apply what I had learned over the previous year while also receiving detailed feedback at each stage. I made mistakes throughout the project, but those mistakes helped me grow. In this post, I’ll walk through each step, explain how I tackled the problem, what I learned, where I struggled, and how the project came together in the end.

Choosing the Topic: From Interest to Research Question

I chose COVID-19 risk prediction for my capstone because COVID-19 was what first got me interested in data visualization. During the pandemic, I remember seeing news reports and dashboards tracking positive cases and realizing how powerful data could be when it was organized clearly. It felt fitting for my capstone to circle back to the topic that started that interest.

For the first part of the project, I focused on choosing a subject, researching the background, and turning a broad idea into clearer research questions. We were taught how to use research tools such as Consensus and Research Rabbit to find relevant previous studies connected to our topic. By the end of this stage, I had a 14-page proposal with a literature review. At the time, I remember wondering how all four homework assignments would eventually fit into one final report without becoming 50 pages long. That was a problem for later, but at this stage, the main goal was learning how to ask the right questions.

My original research questions were weak. They were too long, somewhat unclear, and one of them was really more of a method than a research question. I had to revise them more than once. First, I rewrote them to be clearer and more direct. Later, when putting together the final report, I condensed them even further. These were my final research questions: 

Question 1: Which demographic factors, health behaviors, and pre-existing conditions are most strongly associated with severe COVID-19 outcomes across age groups? 

Question 2: Which combinations of pre-existing conditions and tobacco use are associated with elevated severe outcome risk, and how do these patterns vary by age group? 

Question 3: How do logistic regression and random forest compare in predicting severe COVID-19 outcomes using accuracy, sensitivity, specificity, F1-score, and AUC?

One of the biggest lessons from this stage was that I needed to read the literature more thoroughly and with a more open mind. At first, I was focused heavily on pre-existing conditions and did not fully take in everything the research was showing. Later in the project, I realized I had overlooked an important factor: sex. Since the dataset used “sex” as the variable name, I kept that wording in my analysis. Males showed a higher severe outcome risk, which led me to change direction by including sex in the models and exploring interactions. I also had to go back and reread my sources to make sure the literature review still supported the final direction of the project.

This part of the capstone taught me that choosing a topic is more difficult than I expected. It is not just picking something interesting. The research questions have to be clear, and the research has to be read carefully enough that you are willing to change direction when the evidence points somewhere important.

Data Cleaning Process

For the second part of the capstone, I worked on finding the dataset, cleaning it, handling missing values, and starting exploratory data analysis. I used a Kaggle version of Mexico’s public COVID-19 dataset, which had been partially translated and uploaded for practice and analysis. The original dataset had over 1 million rows, but after filtering to positive COVID-19 cases and cleaning the data, my working dataset had about 388,000 rows and 14 columns.

This dataset had some challenges. One of the biggest issues was that the Kaggle version was not as complete as I originally hoped. The official source included more information, such as patient vitals and additional features, but the Kaggle version was limited to 21 variables. I found the original source, but since it was in Spanish and I did not have enough time to confidently translate and rebuild the dataset from the official version, I stayed with the Kaggle version. If I had more time, I would either work through the official source more carefully or look for another dataset, such as a CDC source, with similar information.

The data also needed to be recoded into a more readable format. Many variables used numeric codes, such as 1 = no and 2 = yes, instead of clear labels. There were also values like 97, 98, and 99, which had to be handled as “other,” “unknown,” or missing depending on the variable. I also engineered my target variable: an adverse outcome, defined as ICU admission, being put on a ventilator, or death.

After cleaning, I removed variables that had almost no variation and would likely add little value to the models to create two datasets to model. From there, I started looking for patterns, investigating possible interaction effects, and creating visualizations to better understand what the data contained.

The EDA was one of the hardest parts of the project for me. Earlier in my degree, I had completed an independent study that included less direct instruction and feedback than a regular course, so I did not feel as prepared for this part of the capstone as I wanted to be. This assignment needed major revision after feedback from my capstone professor, Dr. Cheng. Her notes helped me see what was unclear, and meeting with her made me much more confident in how to approach the EDA.

The biggest lesson I learned was that the EDA is not just about making charts. It is a process. First, the data has to be prepared so the columns make sense, the types are correct, and the values are ready for analysis. Then comes the messy exploration stage, where I looked for patterns, relationships, possible correlations, and anything that might matter later in modeling. Finally, that messy process has to be turned into a clean explanation that other people can understand.

That last step was where I made my biggest mistake. I treated the early exploration too much like the final explanation. In reality, the EDA starts messy, but the report should not stay messy. This part of the capstone taught me that exploring data and communicating data are related, but they are not the same skill. I had to learn how to dig through the data first, then organize the most important findings into something clear, focused, and useful.

Building the Models

This was the main part of the project because it was where I started directly working toward my research questions. I used two supervised learning models, logistic regression (LR) and random forest (RF), to predict adverse COVID-19 outcomes. I also used association rule mining (ARM), an unsupervised data mining method, to look for patterns in the data. The supervised models were focused on prediction, while ARM was more about discovering relationships between variables.

I originally started with the supervised models before moving into ARM. After discussing the project with my professor at the end of the semester, I realized that ARM worked better as part of the exploratory process rather than as a main predictive model. The rules helped point me back toward patterns I needed to think more carefully about, especially interaction effects between sex and pre-existing conditions. Because of that, I rebuilt my LR model to include interactions that I had missed earlier.

Modeling ended up being much more iterative than I expected. I rebuilt the LR model three times as I refined the variables, interactions, and structure. I also learned that hyperparameter tuning for RF can take a lot of time. After reducing the grid, my final search still included 36 models. Each search took over 30 minutes to run, and I had to run it twice each time I ran the report, so computation time became a real constraint.  I ended up reducing the size of the grid to this because the computation time made the original plan less practical.

I also rebuilt parts of the modeling process after receiving feedback from my final presentation. For the final report, I changed the interaction effects in the LR and adjusted thresholds in ARM. These changes gave me results that better captured relationships I had missed before, especially where variables were working together instead of acting alone.

One of my biggest lessons from this stage was that modeling is not a straight path. It is a process of testing, receiving feedback, revising, and sometimes rebuilding from the ground up. I also learned not to choose a method just because it is new and interesting. ARM was useful because it helped support the exploratory side of the project, and I learned a lot from using it, but it was not the strongest tool for the main prediction goal. If I redid the project, I would still use it carefully for pattern discovery, but I would be more intentional about matching each method to the question it is meant to answer.

Learning How to Evaluate Results

After building the models, the next step was evaluating how well they performed on unseen test data. I used metrics such as accuracy, sensitivity, specificity, and confusion matrices to compare LR and RF to each other and to previous research. 

This stage taught me why choosing the right metric matters so much. The adverse outcome class made up only about 15% of the dataset, which meant a model could look accurate by mostly predicting the majority class: no adverse outcome. A null model that predicted every case as no adverse outcome would already look strong by accuracy alone, even though it would fail to identify the cases I actually cared most about.

That is where sensitivity became important. Sensitivity measures how well the model identifies the positive class, which in this project meant cases with an adverse outcome. My LR model did not perform as well as I hoped, with sensitivity of only 21.3%, even after tuning variables and interaction effects with the training data. RF had slightly worse accuracy than the null model, but it had much better sensitivity at 55.7%. Even though RF was not perfect, it was more useful for this problem because it caught more of the adverse outcome cases.

Specificity was also important, but it answered a different question. Specificity measures how well the model identifies the negative class, which in this dataset meant no adverse outcome. This is useful in situations where false positives are especially costly, such as not wanting to accidentally flag someone as guilty when they are innocent. In my project, however, the main concern was identifying higher-risk adverse outcomes, so sensitivity was especially important.

LR and RF were chosen because they were common in the research, explainable, and realistic for the timeline of the project. LR could be explained through odds ratios, while RF could be explained through variable importance plots. Some previous research used other methods, including neural networks and other ensemble models, and part of me wanted to try more supervised models. But with the project timeline and requirement of three methods, I stayed focused on LR, RF, and ARM.

Another lesson was that model performance depends on the data available. My sensitivity was lower than I hoped and lower than some previous research, but one reason was that my dataset did not include clinical values or the fuller set of features available in other studies. I had to evaluate the models in the context of what the dataset could realistically support.

The biggest takeaway from this part of the capstone was simple: choose the right metric for the problem. Accuracy alone can hide important weaknesses, especially with imbalanced data. For this project, sensitivity told a much more meaningful story because it showed whether the models were actually finding the adverse outcome cases.

Turning Assignments into a Final Report

After completing the four homework assignments leading up to the final report, I had about 70 double-spaced pages of work that needed to become one clear, complete report. Instead of trying to stitch everything together, I made a major strategic choice: I started from a blank document and rebuilt the report from scratch, using my homework assignments as references. I wanted the final report to stand alone as a complete project, but I also wanted to apply all the feedback I had received from my professor, the other professors, and my classmates throughout the semester.

This process was about more than just rewriting. I cleaned and refactored my code, organized the project for better reproducibility, rebuilt parts of my ARM process, and revised my models and visuals based on feedback. One of the biggest improvements was in my bar graphs. I changed how I grouped percentages and used fills when comparing different variables with adverse outcomes. The new visuals told the story much more clearly at a quick glance, without making the reader work as hard to understand the comparison.

My focus was to make the final report and code clean, clear, and thorough without making the project longer than it needed to be. The original target was 15–20 pages, and my final report ended up being 30 pages double-spaced. It was longer than the requirement, but the extra length felt necessary to make the analysis readable as a standalone, portfolio-style project rather than just a collection of class assignments.

This part of the capstone taught me that finishing a data science project is not just about running the models. It is also about understanding the data clearly, writing readable code, building visuals that communicate well, and shaping the full analysis into something another person can follow.Stacked bar chart showing the percentage of COVID-19 positive patients by age group, hospitalization status, and adverse outcome. The chart compares patients sent home and hospitalized across young, middle, older, and oldest age groups. 

Before revision, this EDA graph explored the relationship between age group, hospitalization, and adverse COVID-19 outcomes. While the chart included useful information, the percentages were difficult to interpret and the main story was not as clear as it needed to be. 
 
 

This final revision is much clearer than my first version. By improving the age groups and simplifying the visual story, the pattern became easier to understand: older patients were more likely to be hospitalized, and adverse outcomes were especially concentrated among the oldest hospitalized groups.

Final Reflection: What I Learned

This is one of my proudest projects, not just because of the final report, but because of how much I grew through the feedback and revision process. I am grateful for the support I received from my professors and classmates at CMU, because I would not have learned as much without their notes, questions, and guidance.

This is also one reason I started this blog. Now that I am out of school, I want a place to explain my process, reflect on my work, and continue learning from others. I hope to make connections and receive feedback from the data science community, but even without feedback, writing this blog is already helping me think more clearly and grow.

My biggest takeaways from this capstone are that data cleaning matters, models have tradeoffs, and communication is part of data science. Understanding the data and putting it in the correct format is a major part of the work. Bad data leads to bad results. If you miss class imbalance, correlations, outliers, or other important patterns, you may end up backtracking later.

I also learned that every model has tradeoffs. Simpler models are often easier to explain, but they may miss important patterns. More complex models can capture more detail, but they can also become harder to interpret or fit too closely to the training data. The goal is to choose a model that fits the problem: simple enough to explain, but strong enough to capture meaningful patterns in the data.

Just like in teaching, communication is a key part of data science. Results only matter if people can understand what was found, why it matters, and how it can be used. To make findings actionable, the information has to be clearly conveyed not just to technical peers, but also to nontechnical stakeholders.

My next step is to continue what I learned from this capstone and practice the full process again with a new dataset. I have already started my next research-style portfolio project, focused on predicting hospital readmission for diabetic patients. My goal is to work through the full process again, from literature review and EDA to modeling, evaluation, and a final clean report.

If you want to dig into the technical details, you can check out the full report and code here. I’ll keep sharing projects on GitHub, posting updates on LinkedIn, and reflecting here on A Ray of Data. I’d love to hear your thoughts, feedback, or questions in the comments. Thanks for following along as I keep learning and building.

Comments