kafka读取hdfs数据
With the advent of the internet, mobile connectivity, and the Internet of Things (IoT), data collected by organizations and made by each person has grown exponentially. In the last two years alone, over 90% of the world’s data has been created. Every day, 2.5 quintillion bytes of data are produced by humans; 95 million photos and videos are shared on Instagram; 306.4 billion emails are sent, and 5 million Tweets are made.
随着Internet,移动连接和物联网(IoT)的出现,组织收集并由每个人制作的数据呈指数增长。 仅在过去的两年中,已经创建了全球90%以上的数据。 每天,人类会产生2.5万亿字节的数据; Instagram上分享了9500万张照片和视频; 已发送3064亿封电子邮件,并发送了500万条推文。
Big data is becoming business as usual for organizations with the volume, variety, and velocity of data being produced. It is not surprising that companies are looking to create their data strategies by implementing Big Data Technologies to keep up with the surging data and reap the opportunities that data provides.
随着数据量,种类和速度的提高,大数据已成为组织的日常业务。 公司正在寻求通过实施大数据技术来创建自己的数据战略,以跟上不断增长的数据并从数据提供的机会中获得利益,这不足为奇。
What is Big Data Technology?
什么是大数据技术?
Simply put, big data technology is a software-utility used to manage Big Data on a commercial or organization-wide scale. It is designed to analyze, process, and extract the information from extremely complex and large data sets, which the traditional data processing software would never deal with. This technology manages both operational (day-to-day operations) and analytical (business intelligence) big data. Big data technologies are ready to assist in storing huge amounts of data, processing the data, manage data access across teams, and analyze huge data sets.
简而言之,大数据技术是一种用于在商业或组织范围内管理大数据的软件实用程序。 它旨在分析,处理和提取极其复杂的大型数据集中的信息,而传统数据处理软件将无法处理这些信息。 该技术可管理运营(日常运营)和分析(商业智能)大数据。 大数据技术已准备就绪,可以协助存储大量数据,处理数据,管理团队之间的数据访问以及分析庞大的数据集。
In this article, we will compare Hadoop, Kafka, and Data Lake to give a better comparison and understanding of these commonly used big data technologies.
在本文中,我们将比较Hadoop,Kafka和Data Lake,以更好地比较和理解这些常用的大数据技术。
Hadoop
Hadoop的
Hadoop is a framework that allows for the distributed processing of large data sets across clusters of computers using simple programming models. It is designed to scale up from single servers to thousands of machines (1). Hadoop Distributed File System (HDFS) is the storage system of Hadoop which splits big data and distributes across many nodes in a cluster allowing local computation and storage. This also replicates data in a cluster thus providing high availability and backup in case of data loss. Rather than rely on hardware to deliver high-availability, the library itself is designed to detect and handle failures at the application layer, so delivering a highly-available service on top of a cluster of computers, each of which may be prone to failures.
Hadoop是一个框架,它允许使用简单的编程模型在计算机集群之间分布式处理大型数据集。 它旨在从单个服务器扩展到数千台计算机(1)。 Hadoop分布式文件系统(HDFS)是Hadoop的存储系统,它可以拆分大数据并分布在群集中的多个节点上,从而可以进行本地计算和存储。 这还将在集群中复制数据,从而在数据丢失的情况下提供高可用性和备份。 库本身不用于依靠硬件来提供高可用性,而是设计用于检测和处理应用程序层的故障,因此可以在计算机集群的顶部提供高可用性服务,每台计算机都容易出现故障。
Kafka
卡夫卡
Kafka is a distributed system consisting of servers and clients that communicate via a high-performance TCP network protocol (2). It is an open-source software used to process real-time data streams, designed as a distributed transaction log. Simply put, Kafka is a messaging system and an event streaming platform. As same as Hadoop, Kafka replicates topic (similar to a folder in a filesystem), even across geo-regions or datacenters, so that there are always multiple brokers that have a copy of the data just in case things go wrong, you want to do maintenance on the brokers, and so on.
Kafka是一个分布式服务器,由通过高性能TCP网络协议(2)进行通信的服务器和客户端组成。 它是用于处理实时数据流的开源软件,被设计为分布式事务日志。 简而言之, Kafka是一个消息传递系统和一个事件流平台。 与Hadoop一样,Kafka甚至在地理区域或数据中心之间都复制主题(类似于文件系统中的文件夹),因此总是有多个代理具有数据副本,以防万一出错。对经纪人进行维护,等等。
Data Lakes
数据湖
A data lake is a centralized repository that allows the storage of structured and unstructured data at any scale. Data can be stored as-is, without having to first structure the data, and run different types of analytics — from dashboards and visualizations to big data processing, real-time analytics, and machine learning to guide better decisions. A data lake is an architecture in which Hadoop and Kafka are a component of and can run simultaneously or at the same time in the same ecosystem.
数据湖是一个集中的存储库,它可以存储任意规模的结构化和非结构化数据。 数据可以按原样存储,而无需先构造数据并运行不同类型的分析-从仪表板和可视化到大数据处理,实时分析和机器学习,以指导更好的决策。 数据湖是一种架构,其中Hadoop和Kafka是Hadoop的组成部分,可以在同一生态系统中同时或同时运行。
Which one to use?
使用哪一个?
Hadoop, Kafka, and Data Lakes have common functions and might be confusing to differentiate. The table below summarizes the different desired use of big data technology and which software is the most appropriate.
Hadoop,Kafka和Data Lakes具有共同的功能,可能难以区分。 下表总结了大数据技术的不同期望用途,以及哪种软件最合适。
1. Data storage — Hadoop and Data Lakes are most superior and commonly used for large data storage at scale. Although Kafka has a storage option too (via topics), it has better use for connecting and streaming data.
1.数据存储-Hadoop和数据湖是最优越的,通常用于大规模的大型数据存储。 尽管Kafka也有一个存储选项(通过主题),但它可以更好地用于连接和流式传输数据。
2. Data replication — Replication is a smart way to ensure data are backed up and data loss is avoided. Across all technologies, this aspect is present.
2.数据复制-复制是确保备份数据并避免数据丢失的明智方法。 在所有技术中,都存在这一方面。
3. Data/event streaming — Kafka is best for event streaming across multiple platforms (and can be put into the Data Lake). Although Hadoop has sharing capabilities, this occurs mostly in its TCP network, unlike Kafka which can stream data across platforms.
3.数据/事件流-Kafka最适合跨多个平台进行事件流(可以放入Data Lake)。 尽管Hadoop具有共享功能,但这主要发生在其TCP网络中,这与Kafka可以跨平台流数据不同。
4. Data Distribution — all technologies have a default distributed function.
4.数据分发-所有技术均具有默认的分发功能。
5. Data Ecosystem — Data Lakes runs like an ecosystem and a collection of different technology architecture. If a company looks at running a multi-software system, Data Lakes is probably a better
5.数据生态系统-数据湖的运行就像一个生态系统和不同技术架构的集合。 如果公司希望运行多软件系统,Data Lakes可能是一个更好的选择
[1] https://hadoop.apache.org/
[1] https://hadoop.apache.org/
[2] https://kafka.apache.org/intro
[2] https://kafka.apache.org/intro
翻译自: https://medium.com/swlh/a-guide-for-big-data-technology-using-hdfs-kafka-and-data-lake-f9a653f7c1b
kafka读取hdfs数据
相关资源:微信小程序源码-合集6.rar