完全连接的断开连接

    科技2026-09-26  14

    This is the second article in a group of articles intended to help me think through some Deep Learning ideas I’ve been musing about for the last few months. Each article can stand alone, and the first article discussing a simple CNN change is located here .

    这是一组文章中的第二篇文章,旨在帮助我思考过去几个月来我一直在沉思的一些深度学习思想。 每一篇文章都可以独立存在,讨论简单的CNN更改的第一篇文章位于此处。

    For this article I want to explore a fundamental property of the common densely connected neural network and examine a possible change in approach. Be forewarned, these ideas are still very raw and I’m still trying to wrap my mind around an effective way to approach this problem. My hope is that by trying to communicate this on paper the idea(s) will converge to one or more meaningful test cases.

    对于本文,我想探讨常见的紧密连接神经网络的基本属性,并研究方法的可能变化。 请预先警告,这些想法还很原始,我仍在尝试将自己的想法围绕解决该问题的有效方法。 我希望通过尝试在纸上进行交流,想法可以融合到一个或多个有意义的测试用例中。

    一般假设: (General Hypothesis:)

    One of the fundamental building blocks in deep learning is the use of one or more fully connected dense “hidden” layers. This typically takes the form of the primary layer type throughout the network or as the final layer(s) in the more novel neural network architectures. This fully connected layer consists of all the input elements connecting to all of the processing units (aka neurons) in the hidden layer. All of the input elements are then processed by all of the processing units in that layer. My observation of this topology is that although this method is clearly effective it doesn’t appear to be very efficient.

    深度学习的基本组成部分之一是使用一个或多个完全连接的密集“隐藏”层。 在整个网络中,这通常采取主要层类型的形式,或者在更新颖的神经网络体系结构中作为最终层。 此完全连接的层由连接到隐藏层中所有处理单元(即神经元)的所有输入元素组成。 然后,该层中的所有处理单元都将处理所有输入元素。 我对这种拓扑的观察是,尽管这种方法显然有效,但似乎效率不高。

    The thought experiment or analogy I use when thinking about this topology is one of a factory. Imagine each neuron as a factory worker, and that the input element are raw materials. Each worker is provided all of the exact same raw materials as every other worker but each uses a slightly different tool for processing those inputs. Once the outputs from all workers are collected, a determination is made as to what the output most likely should have been. In a sense it is like trying to put Humpty Dumpty back together again, but with each worker acting in isolation with a different tool. Each worker then produces their result and a decider then somehow aggregates the collective output into a reasonable facsimile of Humpty Dumpty. While this analogy has a number of practical flaws, it does provide some insight into why my initial impressions of inefficiency may prove to be valid.

    我在考虑此拓扑时使用的思想实验或类比是工厂之一。 想象每个神经元都是工厂工人,并且输入元素是原材料。 向每个工人提供与每个其他工人完全相同的原材料,但是每个人使用略有不同的工具来处理这些输入。 一旦收集了所有工人的产出,就可以确定最可能的产出是什么。 从某种意义上讲,这就像是尝试再次将“矮胖子”重新放在一起,但每个工人都使用不同的工具孤立地行动。 然后,每个工人产生结果,然后由决策者以某种方式将集体输出汇总为合理的“矮胖”传真。 尽管这种类比有许多实际的缺陷,但它确实提供了一些见识,使我对低效率的最初印象可能被证明是有效的。

    In a physical world, our expectation is that one or more elements of an input are most directly responsible for an outcome, especially when doing things like classification. This follows the same intuition that drives feature engineering as one of the more effective ways to improve consistency and output in a model. Simplistically, the use of an algorithm that requires all input values to be included across multiple decision elements appears to offer at least two areas to explore for efficiency gains.

    在物理世界中,我们期望输入的一个或多个元素最直接地负责结果,尤其是在进行分类之类的操作时。 这遵循了驱动特征工程的直觉,这是提高模型一致性和输出的更有效方法之一。 简而言之,使用要求所有输入值都包含在多个决策元素中的算法,似乎提供了至少两个领域来探索效率提高。

    探索领域: (Areas of Exploration:)

    First, we can look to improve or simplify the input space using techniques similar to those available to the feature engineer. These might include finding ways to reduce the input space size, reduce noise in the input, and improve the scaling and/or distribution of the data. The challenge here is to do each or all of these things in a way that can dynamically adjust while generalizing well across a variety of inputs. It is a form of this capability that may explain the effectiveness of convolutional neural networks and their ability to isolate and emphasize various elements of the input space in ways similar to feature engineering.

    首先,我们可以寻求使用类似于功能工程师可用的技术来改善或简化输入空间。 这些可能包括寻找减少输入空间大小,减少输入中的噪声以及改善数据的缩放和/或分布的方法。 这里的挑战是以一种可以动态调整的方式来完成所有或所有这些事情,同时很好地概括各种输入。 这是这种能力的一种形式,可以解释卷积神经网络的有效性及其以类似于特征工程的方式隔离和强调输入空间的各种元素的能力。

    The other option that follows from the worker analogy is finding a way to subdivide the problem so that inputs are handled by smaller numbers of the processing units(neurons). This points to the possibility of several different relationship patterns between the input element and the processing units that could be explored including:

    从工人的类比得出的另一种选择是找到一种方法来细分问题,以便由较少数量的处理单元(神经元)处理输入。 这表明在输入元素和处理单元之间可能探索几种不同关系模式的可能性,包括:

    one to one: In this case, the algorithm would need to find the single best processing unit to manage the input. In many ways this appears to be more a routing topology or a very simple decision mechanism and would likely be ineffective across complex datasets.

    一对一:在这种情况下,算法将需要找到单个最佳处理单元来管理输入。 在许多方面,这似乎更像是路由拓扑或非常简单的决策机制,并且可能对复杂的数据集无效。

    many to one: The option to be explored here is whether it could be possible to dynamically subsegment the input elements and pass them to the most effective processing unit for that subset of inputs. The goal would be to minimize the processing time while retaining similar levels of efficacy. This would imply that some pre-sorting or filtering algorithm could be trained to sort, filter, and route the input to the corresponding processing unit. This may also imply that the learning could take place in two different arenas. The first targeting how to teach the input processing algorithm to segment and route the inputs effectively with the second targeting the training of the processing unit(neuron) functions in more traditional ways and leveraging large amounts of data.

    多对一:这里要探讨的选项是,是否可以动态细分输入元素并将它们传递给该输入子集的最有效处理单元。 目的是在保持相似水平的功效的同时,将处理时间减至最少。 这意味着可以训练一些预分类或过滤算法以将输入分类,过滤并将其路由到相应的处理单元。 这也可能意味着学习可以在两个不同的领域进行。 第一个目标是如何教输入处理算法有效地分割和路由输入,第二个目标是以更传统的方式训练处理单元(神经元)功能并利用大量数据。

    many to many: This is the existing paradigm and while the all-to-all case has been well explored, there is an opportunity to explore a subclass of the problem that may prove interesting. Similar to the previous case there exists an opportunity to look at algorithms that use groups of samples of the input elements but instead of routing them to individual processing units, instead route them to groups of processing units. These configurations could then be optimized based on the sources subset of data, and the target processing element group.

    多对多:这是现有的范例,尽管对所有案例进行了很好的探讨,但仍有机会探讨可能证明很有趣的问题的子类。 与前面的情况类似,有机会查看使用输入元素样本组的算法,而不是将它们路由到各个处理单元,而是将它们路由到处理单元组。 然后可以根据数据的源子集和目标处理元素组来优化这些配置。

    one to many: This case presents some unique prospects and challenges. Due to the fanout topology, the efficiency could in theory be much worse while at the same time allowing for an increase in output selection. If an expansion could be done dynamically, it might also allow for a different type of learning with some interesting potential. Theoretically, this could enable the dynamic expansion of knowledge and may provide an opportunity to enable redundancy in the learning output. However, to make this real would require rethinking what constitutes an output as the continued expansion would be unlikely to reliably converge into a well-formed output.

    一对多:此案例提出了一些独特的前景和挑战。 由于扇出拓扑,理论上效率可能会差很多,而同时会增加输出选择。 如果可以动态地进行扩展,那么它也可能允许具有某种有趣潜力的另一种类型的学习。 从理论上讲,这可以实现知识的动态扩展,并可以提供在学习输出中实现冗余的机会。 但是,要使其成为现实,就需要重新考虑构成输出的内容,因为持续的扩展不可能可靠地收敛为格式良好的输出。

    测试方法: (Testing Methodologies:)

    Testing this properly will require some additional thought, but my initial ideas are as follows:

    对此进行正确测试将需要一些其他思考,但是我的初步想法如下:

    Focus on the subdivision and routing of the inputs first. Evaluate different methods of selecting and filtering a subset of the input including (but not limited to): Ordered and/or random combination sampling, input hashing by relative value placement (multiple evaluation options could be explored here, as an example relative positioning of highest or lowest values, relative positioning around a key value attribute, structured sampling of fixed size elements across the entire input, and etc.), or some functional selection criteria. Consider treating each input element as a column across the dataset to see if any new insights into how segments might be created could be formed. Look at different column statistics and evaluate the use of the statistics to create column groups.

    首先关注输入的细分和路由。 评估选择和过滤输入子集的不同方法,包括(但不限于):有序和/或随机组合采样,通过相对值放置进行输入哈希(此处可以探索多个评估选项,作为最高位置的相对位置示例)或最低值,围绕键值属性的相对位置,在整个输入中固定大小元素的结构化采样等),或某些功能选择标准。 考虑将每个输入元素视为数据集中的一列,以查看是否可以形成有关如何创建细分的新见解。 查看不同的列统计信息,并评估对创建列组的统计信息的使用。 Evaluate methods for generating target processing groups while minimizing the processing and learning effort. Initial ideas include: locked function neurons, standard learning neurons(activation with backprop), and single neuron stacks with fixed functions(with increasing polynomial degree?). I also want to go back and do a refresher on Fourier transforms as there seems to be some some opportunity to apply those concepts into the creation of a filtering mechanism in this context.

    评估生成目标处理组的方法,同时最大程度地减少处理和学习工作量。 最初的想法包括:锁定函数神经元,标准学习神经元(通过反向传播激活)和具有固定函数的单神经元堆栈(随着多项式度的增加?)。 我也想回过头来对傅立叶变换进行复习,因为在这种情况下似乎有一些机会将这些概念应用到过滤机制的创建中。 Explore the idea of fanout and how this might contribute to expansive learning. Look at the use of fixed and variable size outputs, explore the idea of using multi-directional inputs instead of outputs (this may be a topic for an article by itself), and consider whether simple additive and/or overflow algorithms could be effective.

    探索扇出概念,以及它如何有助于扩展学习。 查看固定大小和可变大小输出的用法,探索使用多方向输入代替输出的想法(这本身可能是一篇文章的主题),并考虑简单的加法和/或溢出算法是否有效。

    下一步: (Next Steps:)

    I’m currently focused on finishing my pursuit of several cloud related certifications, but once that is complete it is my intent to start setting up a set of structured tests on several of these research ideas and provide the results here. Should you happen upon this, and want to chat I’d be happy to have a conversation around the ideas explored here and network on possible avenues of exploration.

    我目前专注于完成对多个与云相关的认证的追求,但是一旦完成,我的目的是开始针对其中的一些研究思路建立一套结构化测试,并在此处提供结果。 如果您碰巧遇到了这个问题,并且想聊天,我很乐意围绕这里探讨的想法进行对话,并就可能的探索途径进行交流。

    翻译自: https://medium.com/@jdchancellor/a-fully-connected-disconnect-6e39ed453763

    Processed: 0.010, SQL: 9