Web news content extraction is vital to improve news indexing and searching in nowadays search engines, especially for the
news searching service. In this paper we study the Web news content extraction problem and propose an automated extraction
algorithm for it. Our method is a hybrid one taking the advantage of both sequence matching and tree matching techniques.
We propose TSReC, a variant of tag sequence representation suitable for both sequence matching and tree matching, along with an associated
algorithm for automated Web news content extraction. By implementing a prototype system for Web news content extraction, the
empirical evaluation is conducted and the result shows that our method is highly effective and efficient.