A US Bike sharing system wants to understand the demand for shared bikes among the people. The company wants to know:
- Which variables are significant in predicting the demand for shared bikes.
- How well those variables describe the bike demands.
- We have to build a model so that management can understand how exactly the demands vary with different features.
The given dataset has around 16 variables and 730 rows. This shows count of registered users(demand) in different weather, on particular day of the week, temperature impacting demand and other factors. This dataset is from year 2018 and 2019.
The assignemnt is divided into below sections:
- Reading and Understanding data
- Visualising the data and performing EDA
- Data preparation(adding dummy variables, diving into train and test data)
- Data modelling using RFE
- Using stats model for analysis for finalizing the model
- Residual Analysis after the model is finalized
- Model evaluation
- Making predictions using final model(model 11 in this case)
- Final Outcome
- Company can expand business during the months of 5,6,7,8 & 9
- Fall season is the most favourable season so company can introduce packages so that more no of bookings can be attracted.
- Company can service the bikes during other months so that business is not impacted.
- Below are the significant factors in determining the demand of bikes during the upcoming year:
- yr
- workingday
- temp
- windspeed
- season winter and summer
- mnth_9
- weekday_6
- weathersit(cloudy, snow, rain)