北京航班延误
Flight delays have become an important subject and problem for air transportation systems all over the world. The aviation industry is continuing to suffer from economic losses associated with flight delays all the time. According to data from the Bureau of Transportation Statistics (BTS) of the United States, more than 20% of U.S. flights were delayed in 2018. These flight delays have a severe economic impact in the U.S. that is equivalent to 40.7 billion dollars per year. Passengers suffer a loss of time, missed business opportunities or leisure activities, and airlines attempting to make up for delays leads to extra fuel consumption and a larger adverse environmental impact. In order to alleviate the negative economic and environmental impacts caused by unexpected flight delays, and balance increasing flight demand with growing flight delays, an accurate prediction of flight delays in airports is needed.
航班延误已成为全世界航空运输系统的重要主题和问题。 航空业一直持续遭受与航班延误有关的经济损失。 根据美国运输统计局(BTS)的数据,2018年美国航班延误了20%以上。这些航班延误对美国造成了严重的经济影响,相当于每年407亿美元。 旅客会浪费时间,错过商机或休闲活动,而航空公司试图弥补延误会导致额外的燃油消耗和更大的不利环境影响。 为了减轻意外的航班延误所造成的负面经济和环境影响,并在不断增长的航班需求与不断增加的航班延误之间取得平衡,需要准确预测机场的航班延误。
Airport delays may result from airlines operations, air traffic congestion, weather, air traffic management initiatives, etc. Most of the reasons are stochastic phenomena which are difficult to predict timely and accurately.
机场延误可能是由航空公司运营,空中交通拥堵,天气,空中交通管理举措等导致的。大多数原因是随机现象,难以及时准确地进行预测。
The goal of this project is to develop a computational model for predicting the delays based on data for flights extracted from Kaggle.
该项目的目标是基于从Kaggle提取的航班数据,开发一种用于预测延误的计算模型。
The first phase is getting data from Kaggle and stores it into PostgreSQL. Second phase is data cleaning. After loading data into database, I cleaned the data mainly depend on business needs. After cleaning all data, next phase is feature engineering, where you create features for machine learning model from raw data. Fourth phase is exploratory data analysis. In this phase I create graphics to understand data. Fifth phase is model analysis, where I applied machine learning algorithms on dataset.
第一阶段是从Kaggle获取数据并将其存储到PostgreSQL中。 第二阶段是数据清理。 将数据加载到数据库后,我清理数据主要取决于业务需求。 清除所有数据之后,下一阶段是功能工程,您可以在其中根据原始数据创建用于机器学习模型的功能。 第四阶段是探索性数据分析。 在此阶段,我将创建图形来理解数据。 第五阶段是模型分析,其中我在数据集上应用了机器学习算法。
All US airlines flights data for 2018 were obtained from Kaggle. As of last count, we have over 7 million rows of on-time performance data stored in a PostgreSQL table that is accessible from jupyter notebook.
美国航空公司2018年的所有航班数据均来自Kaggle。 截至最近一次统计,我们在jupyter Notebook中可以访问的PostgreSQL表中存储了超过700万行的按时性能数据。
This dataset has detail info for airlines, airport, flight number etc. Pretty much all other data is time-related in minutes. It also has the delays broken out by type — like carrier, weather, NAS, security and late aircraft.
该数据集包含有关航空公司,机场,航班号等的详细信息。几乎所有其他数据都与分钟相关。 它还具有按类型细分的延迟-例如航母,天气,NAS,安全和飞机晚点。
At the beginning my dataset had over 7 million flight information. Then I recognize that there were many flights which don’t have all data available. So unavailability of features was the main reason behind eliminating flights from my dataset.
最初,我的数据集包含超过700万个航班信息。 然后,我认识到有很多航班没有所有可用数据。 因此,功能不可用是从数据集中消除飞行的主要原因。
Canceled flights are not delayed flights. If it is canceled that means the flight didn’t happen and values are not helping. I filtered out Canceled flights for my analysis.
取消的航班不是延迟航班。 如果取消,则表示航班没有发生,价值也无济于事。 我筛选出已取消的航班进行分析。
I deleted the null values in the column actual_elasped_time I intend to use in the future.
我删除了我打算在将来使用的actual_elasped_time列中的空值。
After removing those flights, I finally got my dataset with around 6 million data which have all information available.
删除这些航班之后,我最终获得了包含大约600万个数据的数据集,其中包含所有可用信息。
Feature engineering means building additional features out of existing data which is often spread across multiple related tables. Feature engineering requires extracting the relevant information from the data and getting it into a single table which can then be used to train a machine learning model.
特征工程意味着从现有数据中构建附加特征,这些数据通常分布在多个相关表中。 特征工程需要从数据中提取相关信息,并将其放入一个表中,然后该表可用于训练机器学习模型。
Label column “delayed” is created, and value is set to 1 if any delays type columns have a value — like carrier, weather, NAS, security and late aircraft.
创建标签列“ delayed”,如果任何延迟类型列都具有值(例如,承运人,天气,NAS,安全和飞机晚点),则将值设置为1。
Airlines code (op_carrier) and Airport code (origin/dest) columns are converted from “Object” to “int64”.
航空公司代码(op_carrier)和机场代码(origin / dest)列从“对象”转换为“ int64”。
Flight date column fl_date is separated into fl_month, fl_day, and data is updated into these columns from fl_date.
排期日期列fl_date分为fl_month和fl_day,数据从fl_date更新到这些列。
In statistics, exploratory data analysis (EDA) is an approach to analyzing data sets to summarize their main characteristics, often with visual methods. A statistical model can be used or not, but primarily EDA is for seeing what the data can tell us beyond the formal modeling or hypothesis testing task.
在统计中,探索性数据分析(EDA)是一种分析数据集以总结其主要特征的方法,通常使用视觉方法。 可以使用统计模型,也可以不使用统计模型,但是EDA主要用于查看数据可以在形式建模或假设检验任务之外告诉我们的内容。
While there are an almost overwhelming number of methods to use in EDA, one of the most effective starting tools is the pairs plot (also called a scatterplot matrix). A pairs plot allows us to see both distribution of single variables and relationships between two variables. Pair plots are a great method to identify trends for follow-up analysis and, fortunately, are easily implemented in Python.
尽管在EDA中使用了几乎绝大多数方法,但最有效的入门工具之一是结对图(也称为散点图矩阵)。 配对图使我们可以看到单个变量的分布以及两个变量之间的关系。 配对图是识别趋势以进行后续分析的一种很好的方法,幸运的是,可以在Python中轻松实现。
A Heatmap is a graphical representation of data where the individual values contained in a matrix are represented as colors. Heatmaps are perfect for exploring the correlation of features in a dataset. We can now use either Matplotlib or Seaborn to create the heatmap. To get the correlation of the features inside a dataset we can call <dataset>.corr(), which is a Pandas dataframe method. This will give us the correlation matrix.
热图是数据的图形表示,其中矩阵中包含的各个值表示为颜色。 热图非常适合探索数据集中要素的相关性。 现在,我们可以使用Matplotlib或Seaborn来创建热图。 要获取数据<dataset>.corr()的相关性,我们可以调用<dataset>.corr() ,这是Pandas数据<dataset>.corr()方法。 这将给我们相关矩阵。
When we look at the image of the data grouped by the target column delayed, we can see that the data is not distributed well balanced. There are 19% of the values are delayed and 81% of the values are not delayed.
当我们查看由目标列延迟分组的数据的图像时,我们可以看到数据分布不均衡。 有19%的值被延迟,而81%的值没有延迟。
The Machine Learning techniques such as Decision Tree and Logistic Regression have a bias towards the majority class. I used Random Under Sampling algorithm that the majority class has been reduced to the total number of the minority class for handling imbalanced class distribution. Hereby, both classes will have an equal number of entries.
诸如决策树和Logistic回归之类的机器学习技术偏向多数阶级。 我使用随机抽样算法将多数类减少为少数类总数,以处理不平衡的类分布。 因此,两个类将具有相等数量的条目。
In order to run ML algorithms more easily in local, I limited the number of my data to 1 million and decided on my feature columns.
为了在本地更轻松地运行ML算法,我将数据数量限制为一百万,并决定在功能列中使用。
I used 10 algorithms with ensemble models such as Decision Tree, Random Forest, Bagging, KNN etc. I compared train and test accuracy scores, precision, recall, and F1 scores.
我使用了10种算法来集成模型,例如决策树,随机森林,装袋,KNN等。我比较了训练和测试的准确性得分,准确性,召回率和F1得分。
According to the list, the best model was Random Forest. Next, I examined whether I can optimize succeeded models using grid search and randomized search.
根据清单,最好的模型是随机森林。 接下来,我检查了是否可以使用网格搜索和随机搜索来优化成功的模型。
After the parameter estimation, I combined other models to optimized models in a table. I plotted a ROC Curve with the AUC scores below.
在参数估计之后,我将其他模型合并为表格中的优化模型。 我用下面的AUC分数绘制了ROC曲线。
Finally, when we look at the ROC Curve, we can say that Random Forest Classifier will bring us the most accurate results.
最后,当我们查看ROC曲线时,可以说随机森林分类器将为我们带来最准确的结果。
GitHub repository for web scraping and data preprocessing is here.
用于Web抓取和数据预处理的GitHub存储库在此处。
Thank you for your time and reading my article. Please feel free to contact me if you have any questions or would like to share your comments.
感谢您的时间和阅读我的文章。 如果您有任何疑问或想分享您的意见,请随时与我联系。
翻译自: https://medium.com/analytics-vidhya/predicting-a-flight-delays-1582a4238770
北京航班延误
相关资源:离港航班延误的动态预测研究