Flexible protein sequence patterns: a sensitive method to detect weak structural similarities

Geoffrey J. Barton (Lead / Corresponding author), Michael J. E. Sternberg

Research output: Contribution to journalArticlepeer-review

122 Citations (Scopus)


The concept of a flexible protein sequence pattern is defined. In contrast to conventional pattern matching, template or sequence alignment methods, flexible patterns allow residue patterns typical of a complete protein fold to be developed in terms of residue positions (elements), separated by gaps of defined range. An efficient dynamic programming algorithm is presented to enable the best alignment(s) of a pattern with a sequence to be identified. The flexible pattern method is evaluated in detail by reference to the globin protein family, and by comparison to alignment techniques that exploit single sequence, multiple sequence and secondary structural information. A flexible pattern derived from seven globins aligned on structural criteria successfully discriminates all 345 globins from non-globins in the Protein Identification Resource database. Furthermore, a pattern that uses helical regions from just human alpha-haemoglobin identified 337 globins compared to 318 for the best non-pattern global alignment method. Patterns derived from successively fewer, yet more highly conserved positions in a structural alignment of seven globins show that as few as 38 residue positions (25 buried hydrophobic, 4 exposed and 9 others) may be used to uniquely identify the globin fold. The study suggests that flexible patterns gain discriminating power both by discarding regions known to vary within the protein family, and by defining gaps within specific ranges. Flexible patterns therefore provide a convenient and powerful bridge between regular expression pattern matching techniques and more conventional local and global sequence comparison algorithms.

Original languageEnglish
Pages (from-to)389-402
Number of pages14
JournalJournal of Molecular Biology
Issue number2
Publication statusPublished - 20 Mar 1990


  • Amino acid sequence
  • Globins
  • Methods
  • Molecular sequence data
  • Protein conformation
  • Sequence homology, Nucleic Acid


Dive into the research topics of 'Flexible protein sequence patterns: a sensitive method to detect weak structural similarities'. Together they form a unique fingerprint.

Cite this