skip to main content
10.1145/1526709.1526866acmconferencesArticle/Chapter ViewAbstractPublication PagesthewebconfConference Proceedingsconference-collections
poster

Threshold selection for web-page classification with highly skewed class distribution

Published: 20 April 2009 Publication History

Abstract

We propose a novel cost-efficient approach to threshold selection for binary web-page classification problems with imbalanced class distributions. In many binary-classification tasks the distribution of classes is highly skewed. In such problems, using uniform random sampling in constructing sample sets for threshold setting requires large sample sizes in order to include a statistically sufficient number of examples of the minority class. On the other hand, manually labeling examples is expensive and budgetary considerations require that the size of sample sets be limited. These conflicting requirements make threshold selection a challenging problem. Our method of sample-set construction is a novel approach based on stratified sampling, in which manually labeled examples are expanded to reflect the true class distribution of the web-page population. Our experimental results show that using false positive rate as the criterion for threshold setting results in lower-variance threshold estimates than using other widely used accuracy measures such as F1 and precision.

References

[1]
X. He, L. Duan, Y. Zhou and B. Dom, Threshold selection for web-page classification with highly skewed class distribution, Yahoo! Labs Research Report YL-2009-001, 2009
[2]
Y. Yang, A Study on Thresholding Strategies for Text Categorization, Proceedings of SIGIR-01, 24th ACM International Conference on Research and Development in Information Retrieval, 2001

Cited By

View all
  • (2013)What's the deal?Proceedings of the First Australasian Web Conference - Volume 14410.5555/2527208.2527217(69-73)Online publication date: 29-Jan-2013
  • (2012)A feature-free search query classification approach using semantic distanceExpert Systems with Applications: An International Journal10.1016/j.eswa.2012.02.19139:12(10739-10748)Online publication date: 1-Sep-2012
  • (2010)Online stratified samplingProceedings of the 19th ACM international conference on Information and knowledge management10.1145/1871437.1871677(1581-1584)Online publication date: 26-Oct-2010

Index Terms

  1. Threshold selection for web-page classification with highly skewed class distribution

    Recommendations

    Comments

    Information & Contributors

    Information

    Published In

    cover image ACM Conferences
    WWW '09: Proceedings of the 18th international conference on World wide web
    April 2009
    1280 pages
    ISBN:9781605584874
    DOI:10.1145/1526709

    Sponsors

    Publisher

    Association for Computing Machinery

    New York, NY, United States

    Publication History

    Published: 20 April 2009

    Permissions

    Request permissions for this article.

    Check for updates

    Author Tags

    1. binary classifier
    2. skewed class distribution
    3. stratified sampling
    4. threshold selection
    5. web-page classification

    Qualifiers

    • Poster

    Conference

    WWW '09
    Sponsor:

    Acceptance Rates

    Overall Acceptance Rate 1,899 of 8,196 submissions, 23%

    Contributors

    Other Metrics

    Bibliometrics & Citations

    Bibliometrics

    Article Metrics

    • Downloads (Last 12 months)1
    • Downloads (Last 6 weeks)0
    Reflects downloads up to 07 Jan 2025

    Other Metrics

    Citations

    Cited By

    View all
    • (2013)What's the deal?Proceedings of the First Australasian Web Conference - Volume 14410.5555/2527208.2527217(69-73)Online publication date: 29-Jan-2013
    • (2012)A feature-free search query classification approach using semantic distanceExpert Systems with Applications: An International Journal10.1016/j.eswa.2012.02.19139:12(10739-10748)Online publication date: 1-Sep-2012
    • (2010)Online stratified samplingProceedings of the 19th ACM international conference on Information and knowledge management10.1145/1871437.1871677(1581-1584)Online publication date: 26-Oct-2010

    View Options

    Login options

    View options

    PDF

    View or Download as a PDF file.

    PDF

    eReader

    View online with eReader.

    eReader

    Media

    Figures

    Other

    Tables

    Share

    Share

    Share this Publication link

    Share on social media