| |
1
|
I. Bhattacharya and L. Getoor. A latent dirichlet model for unsupervised entity resolution. In Proceedings of the 2006 SIAM International Conference on Data Mining, pages 47--58, 2006.
|
| |
2
|
M. Bilenko and R. Mooney. Adaptive duplicate detection using learnable string similarity measures. In Proceedings of the Ninth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 39--48, 2003.
|
| |
3
|
D. Blei and M. Jordan. Variational inference for dirichlet process mixtures. Bayesian Analysis, 1(1):121--144, 2006.
|
| |
4
|
S.-L. Chuang, K. Chang, and C. Zhai. Context-aware wrapping: Synchronized data extraction. In Proceedings of the Thirty-Third Very Large Databases Conference, pages 699--710, 2007.
|
| |
5
|
V. Crescenzi, G. Mecca, and P. Merialdo. ROADRUNNER: Towards automatic data extraction from large web sites. In Proceedings of the Twenty-Seventh Very Large Databases Conference, pages 109--118, 2001.
|
| |
6
|
J. Ishwaran and L. James. Gibbs sampling methods for stick-breaking priors. Journal of the American Statistical Association, 96(453):161--174, 2001.
|
| |
7
|
J. Lafferty, A. McCallum, and F. Pereira. Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In Proceedings of Eighteenth International Conference on Machine Learning, pages 282--289, 2001.
|
| |
8
|
A. McCallum and D. Jensen. A note on the unification of information extraction and data mining using conditional-probability, relational models. In Proceedings of the IJCAI Workshop on Learning Statistical Models from Relational Data, 2003.
|
| |
9
|
I. Muslea, S. Minton, and C. Knoblock. Hierarchical wrapper induction for semistructured information sources. Journal of Autonomous Agents and Multi-Agent Systems, 4(1-2):93--114, 2001.
|
| |
10
|
K. Probst, M. K. R. Ghai, A. Fano, and Y. Liu. Semi-supervised learning of attribute-value pairs from product descriptions. In Proceedings of the Twentieth International Joint Conference on Artificial Intelligence, pages 2838--2843, 2007.
|
| |
11
|
J. Rurmo, A. Ageno, and N. Catala. Adaptive information extraction. ACM Computing Surveys, 38(2):Article 4, 2006.
|
| |
12
|
S. Sarawagi and W. Cohen. Semi-markov conditional random fields for information extraction. In Advances in Neural Information Processing Systems 17, Neural Information Processing Systems, 2004.
|
| |
13
|
P. Singla and P. Domingos. Entity resolution with markov logic. In Proceedings of the Sixth IEEE International Conference on Data Mining, pages 572--582, 2006.
|
| |
14
|
C. Sutton, K. Rohanimanesh, and A. McCallum. Dynamic conditional random fileds: Factorized probabilistic models for labeling and segmenting sequence data. In Proceedings of Twenty-First International Conference on Machine Learning, pages 783--790, 2004.
|
| |
15
|
Y. Teh, M. Jordan, M. Beal, and D. Blei. Hierarchical dirichlet processes. Journal of the American Statistical Association, 101:1566--1581, 2006.
|
| |
16
|
B. Wellner, A. McCallum, F. Peng, and M. Hay. An integrated, conditional model of information extraction and coreference with application to citation matching. In Proceedings of the 20th Conference on Uncertainty in Artificial Intelligence (UAI), pages 593--601, 2004.
|
| |
17
|
T.-L. Wong and W. Lam. Adapting web information extraction knowledge via mining site invariant and site dependent features. ACM Transactions on Internet Technology, 7(1):Article 6, 2007.
|
| |
18
|
H. Zhao, W. Meng, and C. Yu. Mining templates from search result records of search engines. In Proceedings of the Thirteenth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 884--892, 2007.
|
| |
19
|
J. Zhu, B. Zhang, Z. Nie, J.-R. Wen, and H.-W. Hon. Webpage understanding: an integrated approach. In Proceedings of the Thirteenth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 903--912, 2007.
|