openpose论文

    科技2026-09-25  13

    openpose论文

    介绍(Introduction)

    This paper summary will give you a good understanding of the high-level concept of OpenPose. Since we will be focusing on their creative pipeline and structure, there will be no difficult math or theory included in this summary.

    本文摘要将使您对OpenPose的高级概念有很好的了解。 由于我们将专注于他们的创作渠道和结构,因此本摘要中不会包含困难的数学或理论。

    I would like to start by talking about why I want to share what I learn from this wonderful paper. I have implemented the OpenPose library in my AI Basketball Analysis project. At the time I was building the project, I only knew the basic concept of OpenPose. I spent most of the time working on the code implementation and trying to figure out the best way to combine OpenPose with my original basketball shot detection.

    我首先要谈谈为什么要分享我从这篇精彩论文中学到的知识。 我已经在自己的系统中实现了OpenPose库 AI篮球分析项目。 在我构建项目时,我只知道OpenPose的基本概念。 我花了大部分时间在代码实现上,试图找出将OpenPose与我最初的篮球投篮检测相结合的最佳方法。

    AI Basketball Analysis. Image by Chonyy. AI篮球分析。 图片来自Chonyy。

    Now, as you can see in the GIF, the project is almost completed. I have a full grasp of the implementation of OpenPose after building this project. In order to have a better understanding of what I have been dealing with, I think now it’s time for me to take a deeper look at the research paper.

    现在,正如您在GIF中看到的那样,该项目即将完成。 构建此项目后,我对OpenPose的实现有了充分的了解。 为了更好地了解我所从事的工作,我认为现在是时候深入研究该研究论文了。

    Image taken from “Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields”. 图像取自“使用零件相似性字段进行实时多人二维姿势估计”。

    总览 (Overview)

    In

    在

    The proposed method uses a nonparametric representation, which we refer to as Part Affinity Fields (PAFs), to learn to associate body parts with individuals in the image. This bottom-up system achieves high accuracy and realtime performance, regardless of the number of people in the image.

    所提出的方法使用非参数表示(我们称为部分亲和力字段(PAF))来学习将身体部位与图像中的个体相关联。 无论图像中有多少人,该自下而上的系统都可以实现高精度和实时性能。

    为什么很难? (Why is it difficult?)

    Let’s start by talking about what makes estimating the poses of multi-person in an image so difficult. Here are some difficulties listed.

    让我们从谈论为什么很难估计图像中的多人姿势开始。 这里列出了一些困难。

    Unknown number of people

    人数未知 People can appear at any pose or scale

    人们可以以任何姿势或比例出现People contact and overlapping

    人们接触和重叠Runtime complexity grows with the number of people

    运行时复杂度随着人数的增加而增加

    通用方法(Common Approach)

    OpenPose is definitely not the first team facing this challenge. Then how the other teams try to tackle these problems?

    OpenPose绝对不是面对这一挑战的第一支团队。 那么其他团队如何尝试解决这些问题呢?

    Source: https://medium.com/syncedreview/now-you-see-me-now-you-dont-fooling-a-person-detector-aa100715e396 来源: https : //medium.com/syncedreview/now-you-see-me-now-you-dont-fooling-a-person-detector-aa100715e396

    A

    一种

    This kind of top-down method sounds really intuitive and simple. However, there are some hidden pitfalls in this approach.

    这种自上而下的方法听起来非常直观和简单。 但是,这种方法存在一些隐患。

    Early commitment: no resource to recovery when person detector fails

    早期承诺:当人员检测器发生故障时,没有资源可恢复 Runtime proportional to the number of people

    运行时间与人数成正比Pose estimation is executed even if the person detector fails

    即使人检测器发生故障,也会执行姿势估计

    最初的自下而上的方法(Initial Bottom-Up Approach)

    If the top-down method doesn’t sound like the best approach. Then why don’t we try bottom-up?

    如果自上而下的方法听起来不是最好的方法。 那为什么不尝试自下而上呢?

    Not surprisingly, OpenPose is not the first team that came up with a bottom-up method. Some other teams have also tried the bottom-up approach. However, they are still facing some problems with it.

    毫不奇怪,OpenPose并不是第一个提出自下而上方法的团队。 其他一些团队也尝试了自下而上的方法。 但是,他们仍然面临一些问题。

    Required costly global inference at the final parse

    在最终解析时需要昂贵的全局推断 Didn’t retain the gains in efficiency

    没有保留效率方面的收益Taking several minutes per image

    每个图像花费几分钟

    OpenPose管道(OpenPose Pipeline)

    Overall Pipeline. Image taken from “Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields”. 总体管道。 图像取自“使用零件相似性字段进行实时多人二维姿势估计”。 (a) Take the entire image as the input for a CNN(b) Predict confidence maps for body parts detection(c) Predict PAFs for part association(d) Perform a set of bipartite matching(e) Assemble into a full body pose

    信心图(Confidence Map)

    Confidence map is the 2D representation of the belief that a particular body part can be located​. A single body part will be represented on a single map. So, the number of maps is the same as the total number of the body parts.

    信心图是可以确定特定身体部位的信念的2D表示。 单个身体部位将显示在单个地图上。 因此,地图的数量与身体部位的总数相同。

    The map on the right is only for detecting the left shoulder. Image taken from “Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields”. 右侧的地图仅用于检测左肩。 图像取自“使用零件相似性字段进行实时多人二维姿势估计”。

    PAF (PAF)

    Part Affinity Fields (PAFs), a set of 2D vector fields that encode the location and orientation of limbs over the image domain.

    零件亲和力字段(PAF),一组2D矢量场,可对肢体在图像域上的位置和方向进行编码。

    Image taken from “Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields”. 图像取自“使用零件相似性字段进行实时多人二维姿势估计”。

    双向匹配 (Bipartite Matching)

    Image taken from “Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields”. 图像取自“使用零件相似性字段进行实时多人二维姿势估计”。

    When it comes to finding the full body pose of multiple people, determining Z is a K-dimensional matching problem. This problem is NP-Hard and many relaxations exist. In this work, we add two relaxations to the optimization, specialized to our domain.

    当要找到多个人的全身姿势时,确定Z是一个K维匹配问题。 这个问题是NP-Hard,存在许多松弛。 在这项工作中,我们对优化进行了两次放宽,专门针对我们的领域。

    Relaxation 1: Choose a minimal number of edges to obtain a spanning tree skeleton.

    放松1:选择最少数量的边缘以获得生成树的骨架。 Relaxation 2: Further decompose the matching problem into a set of bipartite matching subproblems. Determine the matching in adjacent tree nodes independently.

    放松二:将匹配问题进一步分解为两部分匹配子问题。 独立确定相邻树节点中的匹配项。

    结构体 (Structure)

    原始结构(Original Structure)

    Image by Chonyy. 图片来自Chonyy。

    The original structure is split into two branches.

    原始结构分为两个分支。

    Beige branch: predicts the confidence map

    米色分支:预测置信度图 Blue branch: predicts the PAF

    蓝枝:预测PAF

    Both branches are organized as an iterative prediction architecture. The predictions from the previous stage are concatenated with the original feature F to produce more refined predictions.

    两个分支都组织为迭代预测体系结构。 来自上一阶段的预测与原始特征F串联在一起,以产生更精确的预测。

    新结构 (New Structure)

    Image taken from “Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields”. 图像取自“使用零件相似性字段进行实时多人二维姿势估计”。

    The first set of stages predicts PAFs L t , while the last set predicts confidence maps S t . The predictions of each stage and their corresponding image features are concatenated for each subsequent stage.

    第一组阶段预测PAF L t,而最后一组阶段预测置信度图S t。 对于每个后续阶段,将每个阶段的预测及其对应的图像特征串联在一起。

    Comparing to their previous publication, they have made a big breakthrough and came up with a new structure. As you can see from the structure above, the confidence map prediction is runned on top of the most refined PAF predictions.

    与以前的出版物相比,他们取得了重大突破并提出了新的结构。 从上面的结构中可以看到,置信度图预测在最精细的PAF预测之上运行。

    Why? The reason is actually really simple.

    为什么? 原因实际上非常简单。

    I

    一世

    那其他图书馆呢?(What about Other Libraries?)

    A growing number of computer vision and machine learning applications require 2D human pose estimation as an input for their systems. The OpenPose’s team is definitely not the only one doing this research. Then why we always think of OpenPose when it comes to pose estimation and not Alpha-Pose? Here are some of the problems with other libraries.

    越来越多的计算机视觉和机器学习应用程序需要2D人体姿势估计作为其系统的输入。 OpenPose的团队绝对不是唯一从事这项研究的人。 那么,为什么在姿势估计而不是Alpha-Pose时总是想到OpenPose? 这是其他库的一些问题。

    Require users to implement most of the pipeline

    要求用户实施大部分管道 Users have to construct their won frame reader

    用户必须构造自己的赢帧阅读器Facial and body keypoint detector are not combined

    面部和身体关键点检测器未组合

    AI篮球分析(AI Basketball Analysis)

    I have implemented OpenPose in this AI Basketball Analysis project. In the beginning, I only have an idea that I want to analyze the shooting pose of the shooter, but I have no clue how to do it! Luckily, I came across OpenPose, which gives me everything I want.

    我已经在此AI Basketball Analysis项目中实现了OpenPose 。 刚开始时,我只有一个想法,就是要分析射击者的射击姿势,但是我不知道该怎么做! 幸运的是,我遇到了OpenPose,它为我提供了我想要的一切。

    Image by Chonyy. 图片来自Chonyy。

    Although the installation process is a little troublesome, the actual code implementation is fairly simple. Their function takes a frame as an input, and output the human coordinate. What’s even better, it could also show the detection overlay on the frame! The simplicity of the code implementation is the main reason why their repo could obatin18.6k+ stars on GitHub.

    尽管安装过程有些麻烦,但是实际的代码实现相当简单。 它们的功能以框架为输入,并输出人体坐标。 更好的是,它还可以在框架上显示检测覆盖图! 代码实现的简单性是最主要的原因,他们的回购可能obatin在GitHub上18.6k +明星。

    My project is below, feel free to check it out!

    我的项目在下面,随时查看!

    翻译自: https://towardsdatascience.com/openpose-research-paper-summary-realtime-multi-person-2d-pose-estimation-3563a4d7e66

    openpose论文

    相关资源:人体姿态估计论文(open pose)
    Processed: 0.008, SQL: 9