递归树的深度

    科技2026-10-03  4

    递归树的深度

    This article is a brief review of the research paper (Socher et al., 2013) in which the authors proposed an efficient, novel approach that focuses on grammatical structure of a sentence for fine-grained sentiment analysis.

    本文是对研究论文( Socher et al。,2013 )的简要回顾,其中作者提出了一种有效,新颖的方法,该方法侧重于句子的语法结构以进行细粒度的情感分析。

    The paper discusses various compositional methods to combine words and phrases (n-gram) to predict the binary (positive or negative) as well as fine-grained (very positive, positive, neutral, negative, very negative) sentiments of words, phrases and whole sentence in a bottom-up fashion. The main contribution of this paper is to introduce a parse tree based dataset with fine-grained sentiment labels: “Stanford Sentiment Treebank” and proposes a neural compositional model: Recursive Neural Tensor Network (RNTN) that outperforms all previous recursive models and achieves state-of-the-art performance.

    本文讨论了将单词和短语(n-gram)组合在一起以预测二进制(正或否定)以及单词(词组和短语的细粒度(非常积极,积极,中立,消极,非常消极)情感)的各种组合方法。自下而上的整个句子。 本文的主要贡献是引入带有细粒度情感标签的基于解析树的数据集:“斯坦福情感树库”,并提出了一种神经组成模型:递归神经张量网络(RNTN),其性能优于以前的所有递归模型并实现了状态-最先进的性能。

    数据集:斯坦福情感树库 (Dataset: Stanford Sentiment Treebank)

    The dataset was created by parsing 11 855 sentences of a movie review excerpt corpus with Stanford Parser, resulting 215,154 phrases which were then randomly sampled and labelled into 25 values ( Figure 1) using Amazon Mechanical Turk. It is observed that shorter phrases have neutral sentiments while more polarized sentiments are being noticed in longer phrases. Also it has been observed based on annotators grading on slider scale, a 5-class classification is enough to capture the major variabilities.

    通过使用Stanford Parser解析电影评论摘录语料库的11 855个句子来创建数据集,从而得到215,154个短语,然后使用Amazon Mechanical Turk将其随机采样并标记为25个值(图1)。 可以看出,较短的短语具有中性的情感,而较长的短语则具有更多的极化情感。 还可以根据基于滑块比例的注释器等级进行观察,采用5级分类足以捕获主要变化。

    The treebank dataset facilitates creating efficient models that can predict the polarity of short sentences and classify difficult negation examples which was not attainable by previous bag-of-words approaches which ignore the word orders in a sentence. Also the binary (positive or negative) classification accuracy on sentiment analysis task crossed 80% mark for the first time after introduction of treebank.

    树库数据集有助于创建有效的模型,该模型可以预测短句的极性并分类困难的否定示例,而这是以前的忽略单词顺序的词袋方法无法实现的。 在引入树库之后,情绪分析任务的二元(正或负)分类准确性也首次超过80%。

    RNTN:递归神经张量网络 (RNTN: Recursive Neural Tensor Network)

    The authors first discusses the compositional methods used by recursive models such as Recursive Neural Network(RNN) and Matrix-Vector Recursive Neural Network (MV-RNN) to predict sentiments of n-gram phrases and their limitations and then proposes RNTN which overcomes the limitations and outperforms all previous models on this task.

    作者首先讨论了递归模型(例如递归神经网络(RNN)和矩阵向量递归神经网络(MV-RNN))用来预测n元语法短语的情感及其局限性的组成方法,然后提出了克服局限性的RNTN并在此任务上胜过所有先前的模型。

    All the recursive models parse the input n-gram into a binary tree, with constituent words as leaves. These words are represented by d-dimensional vectors. The Embedding matrix L ∈ R^{d × |v|} of word vectors (|V | is size of vocabulary) is trained jointly with the models. These word vectors are used to predict the sentiment on word level. Then the recursive models compute the parent vectors ( d-dimensional) in a bottom up fashion ( Figure 2) using various compositional methods, after all of its children vectors are computed. The parent vector at each node is used as input to softmax classifier to compute class probabilities at that node.

    所有的递归模型都将输入的n-gram解析为二叉树,组成词为叶。 这些词由d维向量表示。 与模型一起训练单词向量的嵌入矩阵L∈R ^ {d×| v |}(| V |是词汇量)。 这些单词向量用于预测单词级别的情感。 然后,在计算了所有子向量之后,递归模型使用各种合成方法以自底向上的方式计算父向量(d维)(图2)。 每个节点处的父向量用作softmax分类器的输入,以计算该节点处的类概率。

    Though RNN uses single composition function to compute n-gram vector for phrases at each node, the input vectors interact with each other through a non-linearity ( tanh activation). A more robust and direct interaction between the input vectors is desired, which is achieved in MV-RNN in which each n-gram phrase (n >= 1) is represented by a vector and a matrix. Word vectors and word matrices are the parameters of MV-RNN which are learned during model training, hence with increase in vocabulary size, the number of parameters becomes very large.

    尽管RNN使用单个合成函数为每个节点上的短语计算n元语法向量,但是输入向量通过非线性(tanh激活)相互交互。 在输入向量之间需要更鲁棒和直接的交互,这是在MV-RNN中实现的,其中每个n元语法短语(n> = 1)由向量和矩阵表示。 词向量和词矩阵是在模型训练期间学习的MV-RNN的参数,因此,随着词汇量的增加,参数的数量变得非常大。

    RNTN overcomes these limitation of RNN and MV-RNN; It has lesser number of fixed parameters compared to MV-RNN and uses more powerful and single composition function for all nodes; the input vectors interact explicitly in RNTN unlike standard RNN.

    RNTN克服了RNN和MV-RNN的这些限制; 与MV-RNN相比,它的固定参数数量更少,并且对所有节点使用更强大的单一合成功能; 与标准RNN不同,输入向量在RNTN中进行显式交互。

    见解 (Insights)

    The paper offered several important insights and observations:Models were compared with Naive Bayes, SVMs, BiNB (NB with bigram features), VecAvg(average of word vectors). On fine-grained classification for all phrases (at all node levels of the parse trees) RNTN achieves best performance, followed by MV-RNN, RNN and other models. For binary classification on sentence level, RNTN pushes state of the art accuracy from 80% to 85.4% .

    本文提供了一些重要的见解和观察结果:将模型与朴素贝叶斯,支持向量机,BiNB(具有双字特征的NB),VecAvg(平均单词向量)进行了比较。 在对所有短语(在分析树的所有节点级别)进行细分类时,RNTN表现最佳,其次是MV-RNN,RNN和其他模型。 对于句子级别的二进制分类,RNTN将最新的准确性从80%提高到85.4%。 Optimal performances for all the models were achieved for word vector dimension between 25 and 35, performance deteriorates for smaller and larger value of word vectors which confirms RNTN performance enhancement is not dependent on its increased parameter size as MV-RNN has largest number of parameters.

    对于所有模型,在25至35之间的词向量维数下均实现了最佳性能,当词向量的值越来越大时,性能会下降,这表明RNTN性能增强并不取决于其增加的参数大小,因为MV-RNN具有最多的参数数量。 RNTN reasonably captures the effect of Contrastive Conjunction ( ‘but’ ) on overall sentiment of the sentence.

    RNTN合理地捕捉到了对比性连词(“ but”)对句子整体情感的影响。 RNTN also captures the effect of negation in both positive and negative sentences. It has highest accuracy for negating the positive sentences; it also increases non-negative activation ( degree of non-negative sentiment in a sentence) for negation of negative sentence cases, which clearly indicates the model learns the negation concept well beyond simple negation rules.

    RNTN还可以在正面和负面句子中捕捉到否定的效果。 否定肯定句的准确性最高。 对于否定句子的否定,它也增加了非否定激活(句子中非否定情感的程度),这清楚地表明该模型在简单的否定规则之外学习了否定概念。

    结论 (Conclusions)

    RNTN model is powerful in capturing the structural composition of the words and phrases in a sentence and learning the effect of composition in detecting sentiments in a principled and efficient way. The treebank dataset captures intricacies of linguistic phenomena; all models show substantial improvement on their performances when trained on this new dataset. However, it is to be noted that as RNTN requires the parse tree of the input sentences to be constructed; the model might not perform well in cases of poor grammatical constructions such as dialogues in chatbots or tweets. Another interesting case would be, to observe the effect of pre-trained word embeddings such as word2vec, glove, fasttext on over all performance of the model instead of learning the word vector embeddings as parameters during training.

    RNTN模型在捕获句子中单词和短语的结构组成以及以有原则和有效的方式学习组成在检测情感中的作用方面非常强大。 树库数据集捕获了复杂的语言现象; 在此新数据集上进行训练后,所有模型的性能都得到了显着改善。 但是,需要注意的是,由于RNTN需要构建输入语句的语法分析树; 如果语法结构不佳(例如聊天机器人或推文中的对话),则该模型的效果可能不佳。 另一个有趣的情况是,观察预训练单词嵌入(例如word2vec,g手套,快速文本)对模型的所有性能的影响,而不是在训练过程中学习单词矢量嵌入作为参数。

    翻译自: https://medium.com/@anindyasdas/review-recursive-deep-models-for-semantic-compositionality-over-a-sentiment-treebank-221577eb488

    递归树的深度

    Processed: 0.008, SQL: 9