从头预测

    科技2026-10-08  9

    从头预测

    The data science lifecycle is designed for big data issues and data science projects. Generally, the data science project consists of seven steps which are problem definition, data collection, data preparation, data exploration, data modeling, model evaluation and model deployment. This article goes through the data science lifecycle in order to build a web application for heart disease classification.

    吨他的数据生命周期的科学专为大数据问题和数据科学项目。 通常,数据科学项目包含七个步骤,分别是问题定义,数据收集,数据准备,数据探索,数据建模,模型评估和模型部署。 本文介绍了数据科学生命周期,以便构建用于心脏病分类的Web应用程序。

    If you would like to look at a specific step in the lifecycle, you can reach it without looking deeply to the other steps.

    如果您希望查看生命周期中的特定步骤,则无需深入研究其他步骤就可以实现。

    问题定义 (Problem Definition)

    Clinical decisions are often made based on doctors’ experience and intuition rather than on the knowledge-rich hidden in the data. This leads to errors and much cost that affect the medical services quality. Using analytic tools and data modeling can help in enhancing the clinical decisions. Thus, the goal here is to build a web application to help the doctors in diagnosing heart diseases. The full code of the is available in my GitHub repository.

    临床决策通常基于医生的经验和直觉,而不是基于数据中隐藏的丰富知识。 这会导致错误和大量费用,从而影响医疗服务质量。 使用分析工具和数据建模可以帮助增强临床决策。 因此,这里的目标是建立一个Web应用程序以帮助医生诊断心脏病。 的完整代码可在我的GitHub存储库中找到。

    数据采集 (Data Collection)

    I collected the heart disease dataset from UCI ML. The dataset has the following 14 attributes:

    我从UCI ML收集了心脏病数据集。 数据集具有以下14个属性:

    age: age in years.

    年龄:以年为单位的年龄。

    sex: sex (1=male; 0=female).

    性别:性别(1 =男性; 0 =女性)。

    cp: chest pain type (0 = typical angina; 1 = atypical angina; 2 = non-anginal pain; 3: asymptomatic).

    cp:胸痛类型(0 =典型心绞痛; 1 =非典型心绞痛; 2 =非心绞痛; 3:无症状)。

    trestbps: resting blood pressure in mm Hg on admission to the hospital.

    trestbps:入院时的静息血压,单位为毫米汞柱。

    chol: serum cholesterol in mg/dl.

    chol:血清胆固醇,mg / dl。

    fbs: fasting blood sugar > 120 mg/dl (1=true; 0=false).

    fbs:空腹血糖> 120 mg / dl(1 =真; 0 =假)。

    restecg: resting electrocardiographic results ( 0=normal; 1=having ST-T wave abnormality; 2=probable or definite left ventricular hypertrophy).

    restecg:静息心电图检查结果(0 =正常; 1 = ST-T波异常; 2 =可能或确定的左心室肥大)。

    thalach: maximum heart rate achieved.

    丘脑:达到最大心率。

    exang: exercise induced angina (1=yes; 0=no).

    exang:运动引起的心绞痛(1 =是; 0 =否)。

    oldpeak: ST depression induced by exercise relative to rest.

    oldpeak:运动引起的ST抑郁相对于休息。

    slope: the slope of the peak exercise ST segment (0=upsloping; 1=flat; 2=downsloping).

    坡度:最高运动ST段的坡度(0 =向上倾斜; 1 =平坦; 2 =向下倾斜)。

    ca: number of major vessels (0–3) colored by flourosopy.

    ca:浮游动物上色的主要血管数量(0–3)。

    thal: thalassemia (3=normal; 6=fixed defect; 7=reversable defect).

    thal :地中海贫血(3 =正常; 6 =固定缺损; 7 =可逆缺损)。

    target: heart disease (1=no, 2=yes).

    目标:心脏病(1 =否,2 =是)。

    数据准备与探索 (Data Preparation and Exploration)

    Here is a snapshot of the data header.

    这是数据头的快照。

    The header of the heart disease dataset 心脏病数据集的标题

    From the first look, the dataset contains 14 columns, 5 of them contain numerical values and 9 of them contain categorical values.

    乍看之下,数据集包含14列,其中5列包含数值,其中9列包含分类值。

    The dataset is clean and contains all the information needed for each variable. By using info(), describe(), isnull() functions, no errors, missing values and inconsistencies values are detected.

    数据集是干净的,包含每个变量所需的所有信息。 通过使用info() , describe() , isnull()函数,不会检测到任何错误,缺失值和不一致值。

    #Check null valuesdf.isnull().sum() Null values in the dataset 数据集中的空值

    By checking the percentage of the persons with and without heart diseases, it was found that 56% of the persons in the dataset have heart disease. So, the dataset is relatively balanced.

    通过检查患有和不患有心脏病的人的百分比,发现数据集中56%的人患有心脏病。 因此,数据集是相对平衡的。

    People with and without heart disease in the dataset 数据集中有无心脏病的人

    属性关联(Attributes Correlation)

    This heatmap shows the correlations between the dataset attributes, and how the attributes interact with each other. From the heatmap, we can observe that the chest pain type (cp), exercise induced angina (exang), ST depression induced by exercise relative to rest (oldpeak), the slope of the peak exercise ST segment (slope), number of major vessels (0–3) colored by flourosopy (ca) and thalassemia (thal) are highly correlated with the heart disease (target). We observe also that there is an inverse proportion between the heart disease and maximum heart rate (thalch).

    此热图显示了数据集属性之间的相关性,以及这些属性之间如何相互作用。 从热图中,我们可以观察到胸痛类型(cp),运动引起的心绞痛(exang),运动引起的相对于休息的ST压抑(oldpeak),运动ST段的峰值斜率(slope),主要运动次数萤火虫(ca)和地中海贫血(thal)着色的血管(0–3)与心脏病(目标)高度相关。 我们还观察到,心脏病与最大心率(thalch)之间存在反比。

    Moreover, we can see that the age is correlated with number of major vessels (0–3) colored by flourosopy (ca) and maximum heart rate (thalch). There is also a relation between ST depression induced by exercise relative to rest (oldpeak) and the slope of the peak exercise ST segment (slope). Moreover, there is a relation between the chest pain type (cp) and exercise induced angina (exang). Next, we will analyze these correlations between these features further.

    此外,我们可以看到,年龄与以浮雕(ca)和最大心率(thalch)着色的主要血管数量(0–3)相关。 运动引起的相对于休息的ST压抑(老峰)与峰值运动ST段的斜率(斜率)之间也存在关系。 此外,胸痛类型(cp)与运动诱发的心绞痛(exang)之间存在关系。 接下来,我们将进一步分析这些功能之间的这些相关性。

    1. Age and Maximum Heart Rate

    1.年龄和最大心率

    Heart disease is arising frequently in older people, and the max heart rates are lower for old people with heart disease.

    心脏病在老年人中经常发生,而患有心脏病的老年人的最大心率较低。

    2. Chest Pain

    2.胸痛

    There are four types of chest pain: typical angina, atypical angina, non-anginal pain and asymptomatic. Most of the heart disease patients are found to have asymptomatic chest pain.

    胸痛有四种类型:典型的心绞痛,非典型性心绞痛,非心绞痛和无症状。 发现大多数心脏病患者有无症状的胸痛。

    2. Chest Pain and Exercise Induced Angina

    2.胸痛和运动诱发的心绞痛

    The people who have exercise induced angina; they usually suffer from asymptomatic chest pain, and they are more likely to have heart disease.

    运动引起心绞痛的人; 他们通常患有无症状的胸痛,而且他们更容易患心脏病。

    3. Thalassemia

    3.地中海贫血

    People with reversible defect are likely to have heart disease.

    具有可逆缺陷的人可能患有心脏病。

    4. ST depression and the Slope of the Peak Exercise ST Segment.

    4. ST凹陷和最大运动ST段的斜率。

    The people who have downsloping ST segment have higher values of ST depression and more chance to be infected with heart disease. The greater the ST depression, the greater the chance of disease.

    ST段下坡的人ST抑郁的价值更高,感染心脏病的机会也更多。 ST抑郁症越大,患病的机会越大。

    5. Age and Number of Major Vessels (0–3) Colored by Flourosopy.

    5.主要活动的年龄和数量(0–3),用萤火虫着色。

    Most of the heart disease patients are old and they have one or more major vessels colored by Flourosopy.

    大多数心脏病患者年龄较大,并且有一个或多个由萤石色着色的主要血管。

    资料建模 (Data Modeling)

    Let’s create the machine learning model. We are trying to predict whether a person has heart disease. We will use the ‘target’ column as the class, and all the other columns as features for the model.

    让我们创建机器学习模型。 我们正在尝试预测一个人是否患有心脏病。 我们将' target '列用作类,并将所有其他列用作模型的功能。

    # Initialize data and targettarget = df[‘target’]features = df.drop([‘target’], axis = 1)

    -数据分割 (- Data Splitting)

    We will divide the data into a training set and test set. 80% of the data will be for training and 20% for testing.

    我们将数据分为训练集和测试集。 80%的数据将用于培训,而20%的数据将用于测试。

    # Split the data into training set and testing setX_train, X_test, y_train, y_test = train_test_split(features, target, test_size = 0.2, random_state = 0)

    -机器学习模型 (- Machine Learning Model)

    Here, we will try the below machine learning algorithms then we will select the best one based on its classification report.

    在这里,我们将尝试以下机器学习算法,然后根据其分类报告选择最佳算法。

    Support Vector Machine

    支持向量机 Random Forest

    随机森林Ada Boost

    艾达助推器Gradient Boosting

    梯度提升

    The following function for training and evaluating the classifiers.

    以下功能用于训练和评估分类器。

    def fit_eval_model(model, train_features, y_train, test_features, y_test): results = {} # Train the model model.fit(train_features, y_train) # Test the model train_predicted = model.predict(train_features) test_predicted = model.predict(test_features) # Classification report and Confusion Matrix results[‘classification_report’] = classification_report(y_test, test_predicted) results[‘confusion_matrix’] = confusion_matrix(y_test, test_predicted) return results

    Initialize models, train and evaluate.

    初始化模型,训练和评估。

    # Initialize the modelssv = SVC(random_state = 1)rf = RandomForestClassifier(random_state = 1)ab = AdaBoostClassifier(random_state = 1)gb = GradientBoostingClassifier(random_state = 1)# Fit and evaluate modelsresults = {}for cls in [sv, rf, ab, gb]: cls_name = cls.__class__.__name__ results[cls_name] = {} results[cls_name] = fit_eval_model(cls, X_train, y_train, X_test, y_test)

    Now, we will print the evaluation results.

    现在,我们将打印评估结果。

    # Print classifiers resultsfor result in results: print (result) print()for i in results[result]: print (i, ‘:’) print(results[result][i]) print() print (‘ — — -’) print()

    The results are below:

    结果如下:

    Support Vector Machine Result 支持向量机结果 Random Forest Result 随机森林结果 Ada Boost Results Ada Boost结果 Gradient Boosting Result 梯度提升结果

    From the above results, the best model is Gradient Boosting. So, I will save this model to use it for the web application.

    根据以上结果,最好的模型是Gradient Boosting 。 因此,我将保存此模型以将其用于Web应用程序。

    -保存预测模型 (- Save the Prediction Model)

    Now, we will pickle the model so that it can be saved on disk.

    现在,我们将腌制模型,以便可以将其保存在磁盘上。

    # Save the model as serialized object picklewith open(‘model.pkl’, ‘wb’) as file:pickle.dump(gb, file)

    模型部署 (Model Deployment)

    It is time to start deploying and building the web application using Flask web application framework. For the web app, we have to create:

    现在该开始使用Flask Web应用程序框架部署和构建Web应用程序。 对于Web应用程序,我们必须创建:

    1. Web app python code (API) to load the model, get user input from the HTML template, make the prediction, and return the result.

    1. Web应用程序python代码(API),用于加载模型,从HTML模板获取用户输入,进行预测并返回结果。

    2. An HTML template for the front end to allow the user to input heart disease symptoms of the patient and display if the patient has heart disease or not.

    2.前端HTML模板,允许用户输入患者的心脏病症状并显示患者是否患有心脏病。

    The structure of the files is like the following:

    文件的结构如下:

    /├── model.pkl├── heart_disease_app.py├── templates/ └── Heart Disease Classifier.html

    Web App Python代码 (Web App Python Code)

    You can find the full code of the web app here.

    您可以在此处找到该Web应用程序的完整代码。

    As a first step, we have to import the necessary libraries.

    第一步,我们必须导入必要的库。

    import numpy as npimport picklefrom flask import Flask, request, render_template

    Then, we create app object.

    然后,我们创建app对象。

    # Create applicationapp = Flask(__name__)

    After that, we need to load the saved model model.pkl in the app.

    之后,我们需要将保存的模型model.pkl加载到应用程序中。

    # Load machine learning modelmodel = pickle.load(open(‘model.pkl’, ‘rb’))

    After that home() function is called when the root endpoint ‘/’ is hit. The function redirects to the home page Heart Disease Classifier.html of the website.

    之后,当根端点'/'被命中时,将调用home()函数。 该函数重定向到网站的首页Heart Disease Classifier.html 。

    # Bind home function to URL@app.route(‘/’)def home(): return render_template(‘Heart Disease Classifier.html’)

    Now, create predict() function for the endpoint ‘/predict’. The function is defined this endpoint with POST method. When the user submits the form, the API receives a POST request, the API extracts all data from the form using flask.request.form function. Then, the API uses the model to predict the result. Finally, the function renders the Heart Disease Classifier.html template and returns the result.

    现在,创建predict() 端点'/predict'. 该函数是使用POST方法定义的。 当用户提交表单时,API会收到POST请求,API会使用flask.request.form函数从表单中提取所有数据。 然后,API使用该模型来预测结果。 最后,该函数呈现Heart Disease Classifier.html模板并返回结果。

    # Bind predict function to URL@app.route(‘/predict’, methods =[‘POST’])def predict(): # Put all form entries values in a list features = [float(i) for i in request.form.values()] # Convert features to array array_features = [np.array(features)] # Predict features prediction = model.predict(array_features) output = prediction # Check the output values and retrieve the result with html tag based on the value if output == 1: return render_template(‘Heart Disease Classifier.html’, result = ‘The patient is not likely to have heart disease!’) else: return render_template(‘Heart Disease Classifier.html’, result = ‘The patient is likely to have heart disease!’)

    Finally, start the flask server and run our web page locally on the computer by calling app.run() and then enter http://localhost:5000 on the browser.

    最后,启动flask服务器,并通过调用app.run()在计算机上本地运行我们的网页,然后在浏览器中输入http://localhost:5000 。

    if __name__ == ‘__main__’:#Run the applicationapp.run()

    HTML模板 (HTML Template)

    The following figure presents the HTML form. You can find the code here.

    下图显示了HTML表单。 您可以在此处找到代码。

    The form has 13 inputs for the 13 features and a button. The button sends POST request to the/predict endpoint with the input data. In the form tag, the action attribute calls predict function when the form is submitted.

    该表格有13个功能的13个输入和一个按钮。 该按钮将POST请求与输入数据一起发送到/predict端点。 在form标记中, action属性在提交表单时调用predict功能。

    <form action = “{{url_for(‘predict’)}}” method =”POST” >

    Finally, the HTML page presents the stored result in the result parameter.

    最后,HTML页面将结果存储在result参数中。

    <strong style="color:red">{{result}}</strong>

    概要 (Summary)

    In this article, you learned how to create a web application for prediction from scratch. Firstly, we started with the problem definition and data collection. Then, we worked on data preparation, data exploration, data modeling and model evaluation. Finally, we deployed the model using flask.

    在本文中,您学习了如何创建Web应用程序以从头开始进行预测。 首先,我们从问题定义和数据收集开始。 然后,我们进行了数据准备,数据探索,数据建模和模型评估。 最后,我们使用flask部署了模型。

    Now, it is time to practice and apply what you learn in this article. Define a problem, search for a dataset on the Internet, and then go through the other steps of the data science lifecycle.

    现在,是时候练习并应用本文中学习的内容了。 定义问题,在Internet上搜索数据集,然后执行数据科学生命周期的其他步骤。

    翻译自: https://medium.com/analytics-vidhya/the-lifecycle-to-build-a-web-app-for-prediction-from-scratch-bec1632b5f27

    从头预测

    相关资源:微信小程序源码-合集6.rar
    Processed: 0.124, SQL: 10