递归树的深度
This article is a brief review of the research paper (Socher et al., 2013) in which the authors proposed an efficient, novel approach that focuses on grammatical structure of a sentence for fine-grained sentiment analysis.
本文是对研究论文( Socher et al。,2013 )的简要回顾,其中作者提出了一种有效,新颖的方法,该方法侧重于句子的语法结构以进行细粒度的情感分析。
The paper discusses various compositional methods to combine words and phrases (n-gram) to predict the binary (positive or negative) as well as fine-grained (very positive, positive, neutral, negative, very negative) sentiments of words, phrases and whole sentence in a bottom-up fashion. The main contribution of this paper is to introduce a parse tree based dataset with fine-grained sentiment labels: “Stanford Sentiment Treebank” and proposes a neural compositional model: Recursive Neural Tensor Network (RNTN) that outperforms all previous recursive models and achieves state-of-the-art performance.
本文讨论了将单词和短语(n-gram)组合在一起以预测二进制(正或否定)以及单词(词组和短语的细粒度(非常积极,积极,中立,消极,非常消极)情感)的各种组合方法。自下而上的整个句子。 本文的主要贡献是引入带有细粒度情感标签的基于解析树的数据集:“斯坦福情感树库”,并提出了一种神经组成模型:递归神经张量网络(RNTN),其性能优于以前的所有递归模型并实现了状态-最先进的性能。
The dataset was created by parsing 11 855 sentences of a movie review excerpt corpus with Stanford Parser, resulting 215,154 phrases which were then randomly sampled and labelled into 25 values ( Figure 1) using Amazon Mechanical Turk. It is observed that shorter phrases have neutral sentiments while more polarized sentiments are being noticed in longer phrases. Also it has been observed based on annotators grading on slider scale, a 5-class classification is enough to capture the major variabilities.
通过使用Stanford Parser解析电影评论摘录语料库的11 855个句子来创建数据集,从而得到215,154个短语,然后使用Amazon Mechanical Turk将其随机采样并标记为25个值(图1)。 可以看出,较短的短语具有中性的情感,而较长的短语则具有更多的极化情感。 还可以根据基于滑块比例的注释器等级进行观察,采用5级分类足以捕获主要变化。
The treebank dataset facilitates creating efficient models that can predict the polarity of short sentences and classify difficult negation examples which was not attainable by previous bag-of-words approaches which ignore the word orders in a sentence. Also the binary (positive or negative) classification accuracy on sentiment analysis task crossed 80% mark for the first time after introduction of treebank.
树库数据集有助于创建有效的模型,该模型可以预测短句的极性并分类困难的否定示例,而这是以前的忽略单词顺序的词袋方法无法实现的。 在引入树库之后,情绪分析任务的二元(正或负)分类准确性也首次超过80%。
The authors first discusses the compositional methods used by recursive models such as Recursive Neural Network(RNN) and Matrix-Vector Recursive Neural Network (MV-RNN) to predict sentiments of n-gram phrases and their limitations and then proposes RNTN which overcomes the limitations and outperforms all previous models on this task.
作者首先讨论了递归模型(例如递归神经网络(RNN)和矩阵向量递归神经网络(MV-RNN))用来预测n元语法短语的情感及其局限性的组成方法,然后提出了克服局限性的RNTN并在此任务上胜过所有先前的模型。
All the recursive models parse the input n-gram into a binary tree, with constituent words as leaves. These words are represented by d-dimensional vectors. The Embedding matrix L ∈ R^{d × |v|} of word vectors (|V | is size of vocabulary) is trained jointly with the models. These word vectors are used to predict the sentiment on word level. Then the recursive models compute the parent vectors ( d-dimensional) in a bottom up fashion ( Figure 2) using various compositional methods, after all of its children vectors are computed. The parent vector at each node is used as input to softmax classifier to compute class probabilities at that node.
所有的递归模型都将输入的n-gram解析为二叉树,组成词为叶。 这些词由d维向量表示。 与模型一起训练单词向量的嵌入矩阵L∈R ^ {d×| v |}(| V |是词汇量)。 这些单词向量用于预测单词级别的情感。 然后,在计算了所有子向量之后,递归模型使用各种合成方法以自底向上的方式计算父向量(d维)(图2)。 每个节点处的父向量用作softmax分类器的输入,以计算该节点处的类概率。
Though RNN uses single composition function to compute n-gram vector for phrases at each node, the input vectors interact with each other through a non-linearity ( tanh activation). A more robust and direct interaction between the input vectors is desired, which is achieved in MV-RNN in which each n-gram phrase (n >= 1) is represented by a vector and a matrix. Word vectors and word matrices are the parameters of MV-RNN which are learned during model training, hence with increase in vocabulary size, the number of parameters becomes very large.
尽管RNN使用单个合成函数为每个节点上的短语计算n元语法向量,但是输入向量通过非线性(tanh激活)相互交互。 在输入向量之间需要更鲁棒和直接的交互,这是在MV-RNN中实现的,其中每个n元语法短语(n> = 1)由向量和矩阵表示。 词向量和词矩阵是在模型训练期间学习的MV-RNN的参数,因此,随着词汇量的增加,参数的数量变得非常大。
RNTN overcomes these limitation of RNN and MV-RNN; It has lesser number of fixed parameters compared to MV-RNN and uses more powerful and single composition function for all nodes; the input vectors interact explicitly in RNTN unlike standard RNN.
RNTN克服了RNN和MV-RNN的这些限制; 与MV-RNN相比,它的固定参数数量更少,并且对所有节点使用更强大的单一合成功能; 与标准RNN不同,输入向量在RNTN中进行显式交互。
RNTN model is powerful in capturing the structural composition of the words and phrases in a sentence and learning the effect of composition in detecting sentiments in a principled and efficient way. The treebank dataset captures intricacies of linguistic phenomena; all models show substantial improvement on their performances when trained on this new dataset. However, it is to be noted that as RNTN requires the parse tree of the input sentences to be constructed; the model might not perform well in cases of poor grammatical constructions such as dialogues in chatbots or tweets. Another interesting case would be, to observe the effect of pre-trained word embeddings such as word2vec, glove, fasttext on over all performance of the model instead of learning the word vector embeddings as parameters during training.
RNTN模型在捕获句子中单词和短语的结构组成以及以有原则和有效的方式学习组成在检测情感中的作用方面非常强大。 树库数据集捕获了复杂的语言现象; 在此新数据集上进行训练后,所有模型的性能都得到了显着改善。 但是,需要注意的是,由于RNTN需要构建输入语句的语法分析树; 如果语法结构不佳(例如聊天机器人或推文中的对话),则该模型的效果可能不佳。 另一个有趣的情况是,观察预训练单词嵌入(例如word2vec,g手套,快速文本)对模型的所有性能的影响,而不是在训练过程中学习单词矢量嵌入作为参数。
翻译自: https://medium.com/@anindyasdas/review-recursive-deep-models-for-semantic-compositionality-over-a-sentiment-treebank-221577eb488
递归树的深度
