markwunlp / multiturnresponseselection Goto Github PK

This repo contains our ACL 2017 paper data and source code

Python 100.00%

multiturnresponseselection's Introduction

Douban Conversation Corpus

Data set

We release Douban Conversation Corpus, comprising a training data set, a development set and a test set for retrieval based chatbot. The statistics of Douban Conversation Corpus are shown in the following table.

	Train	Val	Test
session-response pairs	1m	50k	10k
Avg. positive response per session	1	1	1.18
Fless Kappa	N\A	N\A	0.41
Min turn per session	3	3	3
Max ture per session	98	91	45
Average turn per session	6.69	6.75	5.95
Average Word per utterance	18.56	18.50	20.74

The test data contains 1000 dialogue context, and for each context we create 10 responses as candidates. We recruited three labelers to judge if a candidate is a proper response to the session. A proper response means the response can naturally reply to the message given the context. Each pair received three labels and the majority of the labels was taken as the final decision.

As far as we known, this is the first human-labeled test set for retrieval-based chatbots. The entire corpus link https://www.dropbox.com/s/90t0qtji9ow20ca/DoubanConversaionCorpus.zip?dl=0

Data template

label \t conversation utterances (splited by \t) \t response

Source Code

We also release our source code to help others reproduce our result. The code has been tested under Ubuntu 14.04 with python 2.7.

Please first run preprocess.py and edit the code with the correct path, and it will give you a .bin file. After that, please run SMN_Last.py with the generated .bin file, and the training loss will be printed on the screen. If you set the train_flag = False, it will give your predicted score with your model.

Some tips:

The 200-d word embedding is shared at https://1drv.ms/u/s!AtcxwlQuQjw1jF0bjeaKHEUNwitA . The shared file is a list has 3 elements, one of which is a word2vec file. Please Download it and replace the input path (Training data) in my scripy.

Tensorflow resources:

The tensorflow code requires several data set, which has been uploaded on the following path:

Resource file: https://1drv.ms/u/s!AtcxwlQuQjw1jGn5kPzsH03lnG6U

Worddict file: https://1drv.ms/u/s!AtcxwlQuQjw1jGrCjg8liK1wE-N9

Requirement: tensorflow>=1.3

Reference

Please cite our paper if you use the data or code in this repos.

Wu, Yu, et al. "Sequential Matching Network: A New Archtechture for Multi-turn Response Selection in Retrieval-based Chatbots." ACL. 2017.

multiturnresponseselection's People

Stargazers

Watchers

Forkers

iamsile sdsfhtw sunnysai12345 xyz8 hyqleonardo mars-wei little1tow gonewithgt outcastofmusic liulj0507 hydercps hejunqing hellozjj yangqiokay rtygbwwwerr leezqcst zofuthan pollyli cutecha yqf-oo luluxing3 hitxujian seanlee97 speak2me pustar jaffe59 yangliuy xyzhou-puck nonvolatilememory howl-anderson pkulzb kongya skybirdhe skyrunner001 jackylovecoding m13225311263 alphadl minzhang20171225 songyf sungjinlees chutianshen yiwangsun beethovenvirus jankim csypeng tinassh houzhenzhen willie-weichi-yu cyzhangathit andyshenas ubermenschlzy quanpinjie sumhncku st-1997 fendouai snakeztc jt17383 cryptum169 kelvincoco perryhau luomuqinghan xiaoduozhou lc1997 searobbersduck syjbupt ii0 chunlinx henryflee fishewyz dolphin02 minhpqn yorick76ee fengxhao hintonhao bookesse bluecliff mengsuixinhan disenwang cderfdsa dejianyang connietong chatbot-tube henanhorse xingxinyu96 mintsuga fence wdashi-hk agentwx zhangxuemiao chenmoshushi baldr-y aaalex123 chriszhangcx bealeson999 xiaojino xumeng123 daywatch sxhfut mruayan judelee19

multiturnresponseselection's Issues

有点看不懂训练数据

@MarkWuNLP 谢谢
感觉训练数据不应该是一问一答或者多问一答的样子的么？

Evaluation

Input vocabulary size?

If I'm correct the I obtain a vocabulary size larger than 600.000 distinct words from the Ubuntu dataset. That number seems very big. Do you limit the vocabulary size for the model?

Tokenize method

How did you tokenize the raw Douban corpus? I'm trying to test my model with new data, I need to tokenize them firstly.

Error in SMN_Dynamic

I am having an issue with SMN_Dynamic

pygpu.gpuarray.GpuArrayException: ('mismatched shapes', 2)
Apply node that caused the error: GpuDot22(GpuReshape{2}.0, <GpuArrayType<None>(float32, matrix)>)
Toposort index: 756
Inputs types: [GpuArrayType<None>(float32, matrix), GpuArrayType<None>(float32, matrix)]
Inputs shapes: [(2000, 200), (100, 50)]
Inputs strides: [(40000, 4), (200, 4)]
Inputs values: ['not shown', 'not shown']
Outputs clients: [[GpuReshape{3}(GpuDot22.0, MakeVector{dtype='int64'}.0)]]

I'm calling train with 200 for the hidden size and embedding:
train(datasets,wordvecs.W,batch_size=200,max_l=max_word_per_utterence ,hidden_size=200,word_embedding_size=200,model_name='SMN_Dynamic.bin')
Can you help please ?

全部的训练语料能公开么

您好，全部的豆瓣训练语料能对外公开么，谢谢~

Where is the word2vec model for Douban dialogue dataset?

Hi, in your script PreProcess.py, if we use the DoubanConversaionCorpus as you released, where's the coressponding file word2vec.model when we call the class WordVecs?

Where is the training set?

We release Douban Conversation Corpus, comprising a training data set, a development set and a test set

I can only see the test set (test.txt) and a small sample of the training set. Where is the rest of the training (and validation) set?

关于M2的定义

为什么添加矩阵A，而不是像M1一样直接相乘呢？

how to apply douban data to multiturn response model?

hi 您好

如何把豆瓣开源的对话数据应用到这个multiTurnResponseSelection上面呢？
十分感谢

Thanks
weizhen

您好请问下，这个Words词向量设定是要在训练中调整的吗？

https://github.com/MarkWuNLP/MultiTurnResponseSelection/blob/master/src/SMN_Dynamic.py#L95

多谢多谢@MarkWuNLP

关于数据集，是否有未去除标点符号的版本？

您好！感谢您对数据集的公开！关于这份数据集，请问是否有未去除标点符号的版本？

Could you please provide the binary train/test you used in your paper

Hi Wu,

We tried to carefully replicate your work in tensorflow and the performances are awesome, however, our model is still 3-4% lower compared to you reported in your paper.

We suspect the reason is that the word embedding for initialization we used is different to yours. Thus I wonder if you could provide the train/test binary file that you used in your code?

Best wishes,
Xijuan

expecting a more details readme file

How to generate the .pkl file on other datasets

I want to train the model on another dataset, how should I process the data to generate the .pkl files?

format of training file

how is the training file you shared supposed to be read? shall I parse it myself? or have you used a library to store it? I'm talking about this file: https://1drv.ms/u/s!AtcxwlQuQjw1jF0bjeaKHEUNwitA

Tensorflow版本输入错误

在tensorflow版本中，actions从来没有进入到模型训练这一环节。总是把true_utt给response的placeholder。

douban数据

你们的工作对我启发很大，在github上只看到了豆瓣数据的train sample，请问一下完整的数据会发布在哪里，何时发布哦？

麻烦您能提供更详细的数据解读吗？实在找不到代码里数据来源和数据的作用

请问该路径文件如何训练得来

dataset = r"../ubuntu_data.mul.100d.fullw2v.train"

请问下，关于response

response candidate r 我理解应该是固定的一句话，现在是把r这整句话一起输进SMN，

如果换成r的id会怎样？

@MarkWuNLP 谢谢！

SCN.py

URL 失效

onedrive上的URL 失效了，不知道能不能解决一下。感谢

数据集下载链接失效

你好，数据集onedrive下载链接失效了，请问可以回复一下吗？

这个模型预测用的时候，是不是要把库里所有候选都算一遍？

谢谢！
@MarkWuNLP

请问word2vec.model在哪里

您好，按照您在readme里面写的，修改path后运行PreProcess.py提示错误，找不到word2vec.model，请问在PreProcess.py中用到的word2vec.model，以及其他文件：train.topic、mergedic2.txt等文件在哪里呢？谢谢！

弱问下PreProcess.py里word2vec.model文件怎么能有？

这里的
https://github.com/MarkWuNLP/MultiTurnResponseSelection/blob/master/src/PreProcess.py#L149
谢谢谢谢@MarkWuNLP

How many epochs does it take to train UDC?

Thank you for sharing the code for your paper. I was wondering how many epochs you trained the model on Ubuntu dialogue corpus for? Was training done on a single GPU? If so, how long would a single epoch take on average?

请问下您有没考虑过Dialog State Tracking Challenge这个数据集？

@MarkWuNLP 谢谢谢谢

您好，请问下，有没一些，基于比如LSTM生成式的多轮对话方面的paper？

您所了解的情况
多谢
多谢
多谢
@MarkWuNLP

worddict file链接失效了

可以麻烦您重新传一份吗

Questions on running SMN_Last.py / Not work

Hi, Yu,

I followed the instructions in #14 to run SMN_Last.py. But it does not work. As you suggested, I modified the parameter word_embedding_size=200 hidden_size=200 in SMN_Last.py.

The following is the error message. Do you know the reason for this? Thanks!

trainning data 949999 val data 50001
image shape (200, 2, 50, 50)
filter shape (8, 2, 3, 3)
/mnt/scratch/lyang/working/PycharmProjects/NLPIRNNMatchZooQA/src-match-zoo-lyang-dev/matchzoo/conqa/smn_yuwu_acl17/src/CNN.py:233: UserWarning: DEPRECATION: the 'ds' parameter is not going to exist anymore as it is going to be replaced by the parameter 'ws'.
self.output =theano.tensor.signal.pool.pool_2d(input=conv_out_tanh, ds=self.poolsize, ignore_border=True,mode="max")
(200, 50)
[[ -7.18495249e-03 1.15389853e-02 -4.65663650e-03 ..., 1.29073645e-02
5.36817317e-03 -3.73283600e-03]
[ -7.56547710e-04 2.52639169e-03 2.14198747e-03 ..., 1.87613393e-03
3.13413644e-03 -1.80455989e-03]
[ 1.70666858e-02 -1.36987258e-02 2.04090084e-02 ..., 6.40176979e-03
2.26856302e-02 -3.21805640e-02]
...,
[ 1.94880138e-02 1.08157883e-03 7.03368021e-05 ..., 1.80350402e-02
-3.28865572e-03 -1.50815454e-02]
[ 5.96401096e-03 -9.91006641e-04 4.77976783e-03 ..., 2.32200795e-02
3.23144545e-02 -8.65365822e-03]
[ -6.93387402e-04 3.27511476e-03 2.07571058e-03 ..., 1.42423584e-03
3.12527304e-03 8.35354059e-04]]
Traceback (most recent call last):
File "SMN_Last.py", line 422, in
,hidden_size=200,word_embedding_size=200)
File "SMN_Last.py", line 334, in train
train_model = theano.function([index], cost,updates=grad_updates, givens=dic,on_unused_input='ignore')
File "/usr/lib/python2.7/site-packages/theano/compile/function.py", line 326, in function
output_keys=output_keys)
File "/usr/lib/python2.7/site-packages/theano/compile/pfunc.py", line 449, in pfunc
no_default_updates=no_default_updates)
File "/usr/lib/python2.7/site-packages/theano/compile/pfunc.py", line 208, in rebuild_collect_shared
raise TypeError(err_msg, err_sug)
TypeError: ('An update must have the same type as the original shared variable (shared_var=<TensorType(float32, matrix)>, shared_var.type=TensorType(float32, matrix), update_val=Elemwise{add,no_inplace}.0, update_val.type=TensorType(float64, matrix)).', 'If the difference is related to the broadcast pattern, you can call the tensor.unbroadcast(var, axis_to_unbroadcast[, ...]) function to remove broadcastable dimensions.')

豆瓣多轮对话

您好，阅读您的文献的时候发现了您开源的多轮对话语聊，请问能不能把训练集和验证集的语聊发我一下？([email protected]) 非常感谢！

请问依赖的版本都是什么

Hi，我clone下来的代码跑挂了，没弄懂什么原因：

MultiTurnResponseSelection/src/CNN.py:233: UserWarning: DEPRECATION: the 'ds' parameter is not going to exist anymore as it is going to be replaced by the parameter 'ws'.
self.output =theano.tensor.signal.pool.pool_2d(input=conv_out_tanh, ds=self.poolsize, ignore_border=True,mode="max")
Traceback (most recent call last):
File "SMN_Last.py", line 422, in
,hidden_size=100,word_embedding_size=100)
File "SMN_Last.py", line 334, in train
train_model = theano.function([index], cost,updates=grad_updates, givens=dic,on_unused_input='ignore')
File "/usr/lib64/python2.7/site-packages/theano/compile/function.py", line 326, in function
output_keys=output_keys)
File "/usr/lib64/python2.7/site-packages/theano/compile/pfunc.py", line 449, in pfunc
no_default_updates=no_default_updates)
File "/usr/lib64/python2.7/site-packages/theano/compile/pfunc.py", line 208, in rebuild_collect_shared
raise TypeError(err_msg, err_sug)
TypeError: ('An update must have the same type as the original shared variable (shared_var=<TensorType(float32, matrix)>, shared_var.type=TensorType(float32, matrix), update_val=Elemwise{add,no_inplace}.0, update_val.type=TensorType(float64, matrix)).', 'If the difference is related to the broadcast pattern, you can call the tensor.unbroadcast(var, axis_to_unbroadcast[, ...]) function to remove broadcastable dimensions.')

这个是theano的版本问题么？