ABSTRACT
The huge amount of knowledge in web communities has motivated the research interests in threaded discussions. The dynamic nature of threaded discussions poses lots of challenging problems for computer scientists. Although techniques such as semantic models and structural models have been shown to be useful in a number of areas, they are inefficient in understanding threaded discussions due to three reasons: (I) as most of users read existing messages before posting, posts in a discussion thread are temporally dependent on the previous ones; It causes the semantics and structure to be coupled with each other in threaded discussions; (II) in online discussion threads, there are a lot of junk posts which are useless and may disturb content analysis; and (III) it is very hard to judge the quality of a post. In this paper, we propose a sparse coding-based model named SMSS to Simultaneously Model Semantics and Structure of threaded discussions. The model projects each post into a topic space, and approximates each post by a linear combination of previous posts in the same discussion thread. Meanwhile, the model also imposes two sparse constraints to force a sparse post reconstruction in the topic space and a sparse post approximation from previous posts. The sparse properties effectively take into account the characteristics of threaded discussions. Towards the above three problems, we demonstrate the competency of our model in three applications: reconstructing reply structure of threaded discussions, identifying junk posts, and finding experts in a given board/sub-board in web communities. Experimental results show encouraging performance of the proposed SMSS model in all these applications.
- ]]K. Balog, L. Azzopardi, and M. de Rijke. A language modeling framework for expert finding. Information Processing and Management, 06(003):1--12, 2008. Google ScholarDigital Library
- ]]D.M. Blei and J.D. Lafferty. Dynamic topic models. In Proc. of ICML, pages 113--120, 2006. Google ScholarDigital Library
- ]]D.M. Blei, A.Y. Ng, and M.I. Jordan. Latent dirichlet allocation. Journal of Machine Learning Research, 3(6):993--1022, 2003. Google ScholarDigital Library
- ]]W. Buntine and A. Jakulin. Applying discrete PCA in data analysis. In Proc. of UAI, pages 59--66, 2004. Google ScholarDigital Library
- ]]C. Chemudugunta, P. Smyth, and M. Steyvers. Modeling general and specific aspects of documents with a probabilistic topic model. Advances in newral information processing systems, 41(6):391--407, 1990.Google Scholar
- ]]G. Cong, L. Wang, C.-Y. Lin, Y.-I. Song, and Y. Sun. Finding question-answer pairs from online forums. In Proc. 31st SIGIR, pages 467--474, 2008. Google ScholarDigital Library
- ]]S. Ding, G. Cong, C.-Y. Lin, and X. Zhu. Using conditional random Éelds to extract contexts and answers of questions from online forums. In Proc. 11th ACL, pages 710--718, 2008.Google Scholar
- ]]X. Gu and W.-Y. Ma. Building implicit links from content for forum search. In Proc. 29th SIGIR, pages 300--307, 2006. Google ScholarDigital Library
- ]]T. Hofmann. Probabilistic latent semantic indexing. In Proc. 29th SIGIR, pages 50--57, 1999. Google ScholarDigital Library
- ]]J. Huang, M. Zhou, and D. Yang. Extracting chatbot knowledge from online discussion forums. In Proc. 11th IJCAI, pages 423--428, 2006. Google ScholarDigital Library
- ]]J.Scott. Social Network Analysis: A Handbook. Sage Publications, London, 2000.Google Scholar
- ]]J.W. Kim, K.S. Candan, and M.E. Donderler. Topic segmentation of message hierarchies for indexing and navigation support. In Proc. 16th WWW, pages 322--331, 2005. Google ScholarDigital Library
- ]]J. Kleinberg. Authoritative sources in a hyperlinked environment. J. ACM, 46(5):604--622, 1999. Google ScholarDigital Library
- ]]A. McCallum, A. Corrada-Emmanuel, and X. Wang. Topic and role discovery in social networks. In Proc. of IJCAI, pages 249--272, 2007.Google Scholar
- ]]G. Mishne, D. Carmel, and R. Lempel. Blocking blog spam with language model disagreement. In Proc. of AIRWeb, 2005.Google Scholar
- ]]L. Page, S. Brin, R. Motwani, and T. Winograd. The PageRank citation ranking: Bringing order to the web. Technical report. Stanford University, 1998.Google Scholar
- ]]X.-H. Phan, L.-M. Nguyen, and S. Horiguchi. Learning to classify short and sparse text & web with hidden topics from large-scale data collections. In Proc. of WWW, pages 91--100, 2008. Google ScholarDigital Library
- ]]M. Rosen-Zvi, T. Gri±ths, M. Steyvers, and P. Smyth. The author-topic model for authors and documents. In Proc. of UAI, pages 487--494, 2004. Google ScholarDigital Library
- ]]D. Shen, Q. Yang, J.-T. Sun, and Z. Chen. Thread detection in dynamic text message streams. In Proc. 29th SIGIR, pages 35--42, 2006. Google ScholarDigital Library
- ]]X. Song, B.L. Tseng, C.-Y. Lin, and M.-T. Sun. Personalized recommendation driven by information flow. In Proc. 29th SIGIR, pages 509--516, 2006. Google ScholarDigital Library
- ]]C. Wang, D.M. Blei, and D. Heckerman. Continuous time dynamic topic models. In Proc. of UAI, 2008.Google ScholarDigital Library
- ]]J. Zhang, M.S. Ackerman, and L. Adamic. Expertise networks in online communities: structure and algorithms. In Proc. of WWW, pages 221--230, 2007. Google ScholarDigital Library
Recommendations
Modeling semantics and structure of discussion threads
WWW '09: Proceedings of the 18th international conference on World wide webThe abundant knowledge in web communities has motivated the research interests in discussion threads. The dynamic nature of discussion threads poses interesting and challenging problems for computer scientists. Although techniques such as semantic ...
Finding prophets in the blogosphere: bloggers who predicted buzzwords before they become popular
iiWAS '15: Proceedings of the 17th International Conference on Information Integration and Web-based Applications & ServicesIdentifying important users from social media has recently attracted much attention in information and knowledge management community. Although researchers have focused on users' knowledge levels on certain topics or influence degrees on other users in ...
Uses of a private "virtual margin" on public threaded discussions: An exploratory lab-based study
Threaded discussion environments are commonly used to support educational dialogue; however their interfaces do not directly support the private work that students do to interpret and prepare responses to public postings. We examine the use of a private ...
Comments