01 / ContextThe question
behind the work.
Lung-cancer survival varies sharply, but decision-makers need an interpretable view of which patient and disease factors are most strongly associated with two-year outcomes.
My role
I cleaned the SEER cohort, defined survival outcomes, ran statistical tests and logistic regression, exported model outputs, and built the Power BI interpretation layer.
03 / In detailA clinically grounded cohort
I worked with more than 114,000 lung and bronchus cancer cases from the SEER registry between 2004 and 2015. The analysis kept first primary cancers, cleaned survival months, and created two-year survival as the main outcome with five-year survival as a reference view.
The project is personal to me because my father died from cancer. I wanted to use a large public dataset to understand which clinical factors stand out most clearly in the survival record and to build a reporting layer that makes those relationships easier to inspect.
From descriptive tests to a model
I used chi-square tests for stage, sex, and race; t-tests and ANOVA for tumor size; and a multivariable logistic regression across stage, tumor size, positive lymph nodes, sex, and race. Cleaned data and model outputs were exported from R for Power BI.
Stage dominated the results. Localized disease had 8.13 times the modeled odds of two-year survival, regional disease had 3.20 times the odds, and tumor burden and positive lymph nodes moved survival in the opposite direction.
A dashboard for the pattern behind the outcome
The Power BI layer makes the cohort readable through stage, demographic, tumor, and time filters. It pairs survival rates with the regression drivers so the reader can move between raw group patterns and the adjusted model.
Localized cases had a 65.6 percent two-year survival rate compared with 12.0 percent for distant disease. That gap, more than any single dashboard visual, is the central finding of the work.