英国议会上院人工智能报告
Just as 2017 was the year of ransomware and 2018 was the year of crypto mining, 2019 could go down as the year of artificial intelligence (AI). It’s gotten to the point where we’ve had prospects tell us, “If you mention AI, I’m leaving.” So many solutions claim to use AI that organizations are at risk of missing out on beneficial, credible AI-powered products because of the hype and false claims.
就像2017年是勒索软件之年,2018年是加密货币开采之年一样,2019年可能会成为人工智能(AI)年。 到了我们有前景告诉我们的地步,“如果您提到AI,我就要走了。” 如此众多的解决方案声称使用AI,因此,由于炒作和虚假声明,组织有可能错过有益,可靠的AI驱动产品。
Professionals should discern which solutions (and vendors) are credible and which are just jumping on the bandwagon. In this post, I’ll give you some specific questions that will help you tell the difference.
专业人士应该辨别哪些解决方案(和供应商)是可靠的,哪些才是潮流。 在这篇文章中,我将向您提出一些具体的问题,以帮助您区别。
Putting AI Into Perspective
透视AI
But what do I mean by AI exactly? In broad terms, artificial intelligence is all about replicating intelligent human behavior. This often takes the form of machine learning (ML), a subcategory of AI that uses algorithms to categorize data and support predictions. ML functions in settings that are either supervised (mapping input to output variables) or unsupervised (clustering input data with no output variables).
但是AI到底是什么意思? 广义上讲,人工智能就是复制人类的智能行为。 这通常采取机器学习(ML)的形式,这是AI的子类别,它使用算法对数据进行分类并支持预测。 ML在受监督(将输入映射到输出变量)或不受监督(对没有输出变量的输入数据进行聚类)的设置中起作用。
Regardless of the use case, data science (the use of AI, ML, deep learning, etc.) lives and dies by the quality of the data it uses. Take self-driving cars, for example: They use a variety of sensors, such as LiDAR, radar and cameras, to supply an accurate picture of what’s going on. Tests have shown that placing stickers on stop signs can confuse these systems.
无论用例如何,数据科学(使用AI,ML,深度学习等)都会因其使用的数据质量而生存和消亡。 以自动驾驶汽车为例:它们使用各种传感器(例如LiDAR,雷达和摄像头)提供正在发生的事情的准确图像。 测试表明,将贴纸贴在停车标志上会混淆这些系统。
Not all problems with AI are this catastrophic. The AI photo editor FaceApp, which can make a person’s face appear to be older or younger than they are, is making headlines. If AI like this goes wrong, you might get a photo that looks, well, unrealistic. That is not a big problem. In this case, the data itself, and the privacy of app users, could be a problem. Many users will take advantage of AI-driven systems without truly understanding the impact of giving someone else access or ownership to their data. FaceApp applies their age manipulation in the cloud. Those photos could contain sensitive information.
并非AI的所有问题都是灾难性的。 人工智能照片编辑器FaceApp可以使人的面Kong看起来比实际年龄更大或更年轻。 如果像这样的AI出错了,您可能会得到一张看起来很不真实的照片。 那不是大问题。 在这种情况下,数据本身以及应用程序用户的隐私可能会成为问题。 许多用户将无法真正理解授予他人访问或拥有其数据的影响而利用AI驱动的系统。 FaceApp在云端应用他们的年龄控制。 这些照片可能包含敏感信息。
AI isn’t a cure-all for all problems. If done poorly, it can be the source of a whole slew of problems. For instance, in security cases (my company’s industry), AI can flag benign network traffic as malicious, thereby creating false positives. Problems often occur when solutions use a single data source, making them incapable of “connecting the dots” to provide high-fidelity assessments.
人工智能并不是解决所有问题的万灵药。 如果做得不好,那可能是一大堆问题的根源。 例如,在安全情况下(我公司的行业),AI可以将良性网络流量标记为恶意,从而产生误报。 当解决方案使用单个数据源时,通常会出现问题,使其无法“连接点”来提供高保真评估。
Given these potential drawbacks, organizations can get the most bang for their buck by being discerning about which AI solution they want to purchase. But where and how do organizations start when they’re evaluating a potential tool, regardless of the use case? I hope these six questions will help.
鉴于这些潜在的缺点,组织可以通过了解要购买的AI解决方案来最大程度地发挥自己的优势。 但是,无论用例如何,组织在评估潜在工具时从何处开始?如何开始? 我希望这六个问题会有所帮助。
Question 1: What types of AI do you use?
问题1:您使用哪种类型的AI?
Weak answer: The vendor provides only one type. This indicates that they’re looking at just one problem. If they’re looking at a problem in only one way, they’re likely missing what they could learn from other techniques. It’s like trying to solve a crime by talking to only one of the 10 suspects and witnesses or considering only one type of evidence. You don’t get the full story.
回答不正确:供应商仅提供一种类型。 这表明他们只关注一个问题。 如果他们仅以一种方式看问题,那么他们很可能会错过从其他技术中学到的东西。 这就像通过与10名犯罪嫌疑人和证人中的仅一名交谈或仅考虑一种证据来解决犯罪一样。 您没有完整的故事。
Strong answer: The solution provider uses more than one type of AI. The vendor should be able to give you a clear answer on where and why they use AI. It’s not just about using more than one type but using AI in different aspects of their analysis. Even using the same type of AI in multiple locations is better than a single type in one location. Ideally, these types include various combinations of AI, each of which is particularly well-suited to analyze different types of data or answer different types of questions.
强有力的答案:解决方案提供商使用多种类型的AI。 供应商应该能够为您提供明确的答案,说明他们在何处以及为何使用AI。 这不仅涉及使用一种以上的类型,还涉及在分析的不同方面使用AI。 即使在多个位置使用相同类型的AI也比在一个位置使用单个类型的AI更好。 理想情况下,这些类型包括AI的各种组合,每种组合特别适合分析不同类型的数据或回答不同类型的问题。
Question 2: What are the specific algorithms you use in your ML?
问题2:您在ML中使用的具体算法是什么?
Weak answer: The vendor can’t name any. This lack of familiarity could illustrate a general lack of knowledge about AI. It also typically means that whatever AI algorithms the solutions provider is using are likely not configured in a way that will add maximum value.
答案很微弱:供应商无法命名。 这种不熟悉可能说明普遍缺乏有关AI的知识。 这通常还意味着,解决方案提供商使用的任何AI算法都可能不会以增加最大值的方式进行配置。
Strong answer: The solutions provider names specific and varied algorithms that it has incorporated into its tools and explains why those algorithms are appropriate and beneficial. These may include the following:
强有力的答案:解决方案提供商列出了已整合到其工具中的特定且多样化的算法,并解释了为什么这些算法合适且有益。 这些可能包括以下内容:
* DBScan: This is useful in unsupervised ML.
* DBScan:这在无监督的ML中很有用。
* Isolation Forest: Data scientists generally use this algorithm for supervised ML.
*隔离林:数据科学家一般都采用这种算法进行监督ML。
Naïve Bayes, Lasso Regression, LSH/Minhash — the list goes on and on. The point is that the solutions provider should be able to name at least some of these and their objectives. For example, if you need new tires, it’s not enough to blindly buy all-weather tires. You need to understand why you need all-weather tires to know that they’re the appropriate product.
朴素的贝叶斯,套索回归,LSH / Minhash-清单还在不断增加。 关键是,解决方案提供商应至少能够列举其中一些及其目标。 例如,如果您需要新轮胎,仅仅盲目购买全天候轮胎是不够的。 您需要了解为什么需要全天候轮胎才能知道它们是合适的产品。
Question 3: How are your ML models generated, trained, scored and validated, and who does this?
问题3:您的ML模型是如何生成,训练,评分和验证的,这是谁做的?
Weak answer: If the vendor cannot provide some information about this aspect of using AI, then it’s highly likely they’re either very new at it or don’t understand how to utilize it effectively. Training, scoring and validating models is a critically important exercise if there’s any hope that the algorithms will deliver as expected. And it’s not something that just anyone can do.
答案很弱:如果供应商无法提供有关使用AI方面的某些信息,那么很有可能他们是AI的新手,或者不了解如何有效地利用AI。 如果有希望算法能够按预期提供的话,训练,评分和验证模型是至关重要的工作。 这不是任何人都能做的。
Strong answer: The solutions provider says it’s different for supervised and unsupervised applications. Also, listen for references to the data sets that each algorithm models. Different file types (like .exe, .pdf and macros) and network traffic (like HTTP, DNS, SMTP) all have separate training sets, so it’s useful to evaluate all these resources individually. Think of it as a decathlete who has different training regimens for each of the events — the 1,500-meter, the pole vault, the javelin and so on. They should specifically be using some kind of scoring system that they can use to inform users about what they should ultimately be doing with the models’ results. And this type of work is typically done by a data scientist.
强有力的答案:该解决方案提供商表示,对于受监督和不受监督的应用程序,情况有所不同。 另外,请侦听对每种算法建模的数据集的引用。 不同的文件类型(例如.exe,.pdf和宏)和网络流量(例如HTTP,DNS,SMTP)都有单独的训练集,因此分别评估所有这些资源非常有用。 可以将它想象成一个十项全能运动员,他对每项赛事都有不同的训练方式-1,500米,撑竿跳高,标枪等等。 他们应该特别使用某种评分系统,可以用来告知用户他们最终将如何处理模型结果。 这种工作通常由数据科学家完成。
In the next part of this series, we’ll cover three more questions: What are the goals of each algorithm, how often are your ML models updated, and what happens when your ML makes a bad decision?
在本系列的下一部分中,我们将再讨论三个问题:每种算法的目标是什么,您的ML模型多久更新一次,并且当ML做出错误的决定时会发生什么?
Originally published at https://www.forbes.com.
最初在https://www.forbes.com上发布。
翻译自: https://medium.com/@brian.laing/council-post-three-questions-to-ask-ai-solution-vendors-that-claim-they-use-artificial-51152f02d3c0
英国议会上院人工智能报告
相关资源:四史答题软件安装包exe