Back to results
Bibliographic record · Consultation and access
Artículo

Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences

Weizhong Li; Adam Godzik · Bioinformatics · 2006

Resource page
Quick overview. Review the resource’s basic details, then access the content using the main button. This page shows only the information needed to identify, cite, and open the work.

Resource access

Open the content from the main option or choose another available source.

OpenAlex OpenAlex Works
Entrar por OpenAlex
Main access

Resource page

Resource reference page. Full text availability has not been automatically confirmed.
Open resource

Summary

Descripción general del contenido del recurso.

MOTIVATION: In 2001 and 2002, we published two papers (Bioinformatics, 17, 282-283, Bioinformatics, 18, 77-82) describing an ultrafast protein sequence clustering program called cd-hit. This program can efficiently cluster a huge protein database with millions of sequences. However, the applications of the underlying algorithm are not limited to only protein sequences clustering, here we present several new programs using the same algorithm including cd-hit-2d, cd-hit-est and cd-hit-est-2d. Cd-hit-2d compares two protein datasets and reports similar matches between them; cd-hit-est clusters a DNA/RNA sequence database and cd-hit-est-2d compares two nucleotide datasets. All these programs can handle huge datasets with millions of sequences and can be hundreds of times faster than methods based on the popular sequence comparison and database search tools, such as BLAST.

How to cite

Elegí el formato que necesitás y copiá la referencia al portapapeles.

APA 7

Li, W. & Godzik, A. (2006). Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences. https://doi.org/10.1093/bioinformatics/btl158

MLA

Li, Weizhong, and Adam Godzik. "Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences." 2006. https://doi.org/10.1093/bioinformatics/btl158.

Chicago

Li, Weizhong and Adam Godzik. 2006. "Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences.". https://doi.org/10.1093/bioinformatics/btl158.

Harvard

Li, W. and Godzik, A. 2006, Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences, Bioinformatics, available at: https://doi.org/10.1093/bioinformatics/btl158 [Accessed 8 Aug. 2026].

Share and print

Save the record, copy its permanent link, or print it as a PDF.

Export reference

You can export the record in common formats for use in a reference manager.

Resource details

Bibliographic information to help confirm that this is the correct material.

Title
Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences
Author / contributors
Weizhong Li; Adam Godzik
Publisher
Bioinformatics
Publication year
2006
Language
English

Subjects

Explore related resources through these subjects.

Copied