NAME: DATE: Lab06 – Literature summary on Homology Modeling (50 pts) The purpose of this assignment is to summarize and evaluate existing "literature" (or published material) in order to establish current knowledge on the subject of

profileMichelle_Michy
rkjacobson_annrepmedchem_2004.pdf

Comparative Protein Structure Modeling and its

Applications to Drug Discovery

Matthew Jacobson 1

and Andrej Sali 1,2

1 Department of Pharmaceutical Chemistry, California Institute for

Quantitative Biomedical Research, Mission Bay Genentech Hall, 600 16th Street,

University of California, San Francisco, CA 94143-2240, USA 2 Department of Biopharmaceutial Sciences, California Institute for

Quantitative Biomedical Research, Mission Bay Genentech Hall, 600 16th Street,

University of California, San Francisco, CA 94143-2240, USA

Contents

1. Introduction 259

2. Fold assignment and sequence-structure alignment 261

3. Comparative model building 261

4. Loop modeling 262

5. Sidechain modeling 263

6. Comparative modeling by MODELLER 264

7. Physics-based approaches to comparative model construction and refinement 264

8. Accuracy of comparative models 266

9. Modeling on a genomic scale 266

10. Applications of comparative modeling to drug discovery 267

10.1. Comparative models vs experimental structures in virtual screening 267

10.2. Use of comparative models to obtain novel drug leads 268

10.3. Comparative models of kinases in virtual screening 269

10.4. GPCR comparative models for drug development 270

10.5. Other uses of comparative models in drug development 271

10.6. Future directions 272

11. Conclusions 273

References 273

1. INTRODUCTION

Homology or comparative protein structure modeling constructs a three-dimensional

model of a given protein sequence based on its similarity to one or more known

structures. In this perspective, we begin by describing the comparative modeling

technique and the accuracy of the models. We then discuss the significant role that

comparative prediction plays in drug discovery. We focus on virtual ligand screening

against comparative models and illustrate the state-of-the-art by a number of specific

examples.

The genome sequencing efforts are providing us with complete genetic blueprints for

hundreds of organisms, including humans. We are now faced with describing,

ANNUAL REPORTS IN MEDICINAL CHEMISTRY, VOLUME 39 q 2004 Elsevier Inc. ISSN: 0065-7743 DOI 10.1016/S0065-7743(04)39020-2 All rights reserved

controlling, and modifying the functions of proteins encoded by these genomes. This

task is generally facilitated by protein three-dimensional structures [1], which are best

determined by experimental methods such as X-ray crystallography and nuclear

magnetic resonance (NMR) spectroscopy. Despite significant advances in these

techniques, many protein sequences are not easily accessible to structure determination

by experiment. Over the last two years, the number of sequences in the comprehensive

public sequence databases, such as SwissProt/TrEMBL [2] and GenPept [3], increased

by a factor of 2.3 from 522,959 to 1,215,803 on 26 April 2004. In contrast, despite

structural genomics, the number of experimentally determined structures deposited in

the Protein Data Bank (PDB) increased by only a factor of 1.4 over the same period,

from 16,612 to 23,793 [4]. Thus, the gap between the numbers of known sequences and

structures continues to grow.

Protein structure prediction methods attempt to bridge this gap [5]. The first class of

protein structure prediction methods, including threading and comparative modeling, rely

on detectable similarity spanning most of the modeled sequence and at least one known

structure [6]. The second class of methods, de novo or ab initio methods, predict the

structure from sequence alone, without relying on similarity at the fold level between the

modeled sequence and any of the known structures [7]. Despite progress in ab initio

protein structure prediction [7,8], comparative modeling remains the most reliable

method to predict the 3D structure of a protein with an accuracy that can be comparable to

a low-resolution, experimentally determined structure [6].

Comparative modeling is carried out in four sequential steps: finding known structures

(templates) related to the sequence to be modeled (target), aligning the target sequence

with the templates, building the model, and assessing the model. Therefore, comparative

modeling is only applicable when the target sequence is detectably related to a known

protein structure. Using automated comparative modeling, the fraction of sequences with

comparative models for at least one domain remained at ,57% over the last two years [9]. A number of servers for automated comparative modeling are available (http://

salilab.org/bioinformatics_resources.shtml). Automation makes comparative modeling

accessible to both experts and the non-specialists alike. Many of the servers are tested at

the bi-annual CAFASP meetings and continually by the LiveBench and EVA [10 – 14]

web servers for assessment of automated structure prediction methods. However, in

spite of automation, manual intervention is generally still needed to maximize the

accuracy of the models.

Comparative modeling will benefit from structural genomics [15]. Structural genomics

aims to structurally characterize most protein sequences by an efficient combination of

experiment and prediction [16,17]. This aim will be achieved by careful selection of

target proteins and their structure determination by X-ray crystallography or NMR

spectroscopy. There are a variety of target selection schemes, ranging from focusing on

only novel folds to selecting all proteins in a model genome [1]. A model-centric view

requires that targets be selected such that most of the remaining sequences can be

modeled with useful accuracy by comparative modeling. Even with structural genomics,

the structure of most proteins will be modeled, not determined by experiment.

As discussed below, the accuracy of comparative models, and correspondingly the

variety of their applications, decreases sharply below the 30% sequence identity cutoff,

mainly as a result of a rapid increase in alignment errors. Thus, structural genomics aims

M. Jacobson and A. Sali260

to determine protein structures such that most of the remaining sequences are related to at

least one known structure at higher than 30% sequence identity [1,15,17]. It was recently

estimated that this cutoff requires a minimum of 16,000 structures to cover 90% of all

protein domain families, including those of membrane proteins [17]. These 16,000

structures will allow the modeling of a very much larger number of proteins. For

example, the New York Structural Genomics Research Consortium measured the impact

of its structures by documenting the number and quality of the corresponding models for

detectably related proteins in the non-redundant sequence database. For each new

structure, on average ,100 protein sequences without any prior structural characteriz- ation could be modeled at least at the fold level [9]. This large leverage of structure

determination by protein structure modeling illustrates and justifies the premise of

structural genomics.

We begin by reviewing the methods needed for each of the four steps of comparative

modeling. While we only briefly touch on fold assignment, sequence-structure alignment,

and model assessment, we elaborate on model building; we emphasize the modeling of

loops and sidechains, because of their importance in ligand docking and rational drug

discovery. We continue by describing MODBASE, a comprehensive database of

comparative models. Next, we describe the role of comparative modeling in drug

discovery, focusing on ligand docking against comparative models. We compare

successes of docking against models and x-ray structures, and illustrate the computational

docking against models with a number of examples, including kinases and G-protein

coupled receptors.

2. FOLD ASSIGNMENT AND SEQUENCE-STRUCTURE ALIGNMENT

The templates for modeling may be found by pairwise sequence alignment methods, such

as BLAST and FASTA, profile-sequence alignment methods, such as PSI-BLAST,

profile – profile alignment methods, such as SALIGN, Hidden Markov Models, such as

SAM-T02, and sequence-structure threading methods that can sometimes reveal more

distant relationships than purely sequence-based methods [18 – 23]. Threading methods

assign the fold by threading the sequence through each of the structures in a library of all

known folds; each sequence-structure alignment is assessed by the energy of a

corresponding coarse model, not by sequence similarity as in sequence comparison

methods. Recently, the accuracy of aligning a sequence to a remotely related protein

structure has been improved by a genetic algorithm protocol that iterates through

alignment, model building, and model assessment [24].

3. COMPARATIVE MODEL BUILDING

Comparative protein structure prediction produces an all-atom model of a sequence,

based on its alignment to one or more related protein structures. Comparative model

building includes either sequential or simultaneous modeling of the core of the protein,

loops, and side-chains. In the original comparative approach, a model is constructed from

Comparative Protein Structure Modeling and its Applications to Drug Discovery 261

a few template core regions, and from loops and side-chains obtained from either aligned

or unrelated structures [25 – 27]. Another family of comparative methods relies on

approximate positions of conserved atoms from the templates to calculate the coordinates

of other atoms [28]. A third group of methods uses either distance geometry or

optimization techniques to satisfy spatial restraints obtained from the sequence-template

alignment [29 – 31]. Next, we review a large variety of specialized methods that focus on

the modeling of loops in a fixed environment of the rest of the protein and the modeling of

sidechains on a fixed backbone.

4. LOOP MODELING

In comparative modeling, target sequences often have residues inserted relative to the

template structures or have regions that are structurally different from the corresponding

regions in the templates. Thus, no structural information about these segments can be

extracted from the template structures. These regions frequently correspond to surface

loops. Loops often play an important role in defining the functional specificity of a given

protein, forming the active and binding sites. The accuracy of loop modeling can be a

major factor determining the usefulness of comparative models in applications such as

ligand docking. Loop modeling can be seen as a mini protein folding problem because the

correct conformation of a given segment of a polypeptide chain has to be calculated

mainly from the sequence of the segment itself. However, loops are generally too short to

provide sufficient information about their local fold. Even identical decapeptides in

different proteins do not always have the same conformation [32,33]. Some additional

restraints are provided by the core anchor regions that span the loop and by the structure

of the rest of a protein that cradles the loop. Although many loop modeling methods have

been described, it is still challenging to model correctly and confidently loops longer than

approximately 8 – 10 residues [34,35].

There are two main classes of loop modeling methods: (i) database search

approaches that scan a database of all known protein structures to find segments fitting

the anchor core regions [36,37]; (ii) conformational search approaches that rely on

optimizing a scoring function [38 – 40]. There are also methods that combine these two

approaches [41,42].

The database search approach to loop modeling is accurate and efficient when a

database of specific loops is created to address the modeling of the same class of loops,

such as b-hairpins [43], or loops on a specific fold, such as the hypervariable regions in

the immunoglobulin fold [37,44]. There are attempts to classify loop conformations into

more general categories, thus extending the applicability of the database search approach

[45 – 47]. However, the database methods are limited because the number of possible

conformations increases exponentially with the length of a loop. As a result, only loops

up to 4 – 7 residues long have most of their conceivable conformations present in the

database of known protein structures [48,49]. This limitation is made even worse by

the requirement for an overlap of at least one residue between the database fragment and

the anchor core regions, which means that modeling a five residue insertion requires at

least a seven residue fragment from the database [50]. Despite the rapid growth

of the database of known structures, it does not seem possible to cover most

M. Jacobson and A. Sali262

of the conformations of a 9-residue segment in the foreseeable future. On the other hand,

most of the insertions in a family of homologous proteins are shorter than 10 – 12

residues [34].

To overcome the limitations of the database search methods, conformational search

methods were developed [38,39]. There are many such methods, exploiting different

protein representations, objective functions, and optimization or enumeration algorithms.

The search algorithms include the minimum perturbation method, molecular dynamics

simulations, genetic algorithms, Monte Carlo and simulated annealing, multiple copy

simultaneous search, self-consistent field optimization, and enumeration based on graph

theory [41,51 – 59]. The accuracy of loop predictions can be further improved by

clustering the sampled loop conformations and partially accounting for the entropic

contribution to the free energy [60]. Another way to improve the accuracy of loop

predictions is to consider the solvent effects. Improvements in implicit solvation models,

such as the Generalized Born solvation model, motivated their use in loop modeling. The

solvent contribution to the free energy can be added to the scoring function for

optimization, or it can be used to rank the sampled loop conformations after they are

generated with a scoring function that does not include the solvent terms [34,61 – 64].

5. SIDECHAIN MODELING

Two simplifications are frequently applied in the modeling of sidechain conformations.

First, amino acid residue replacements often leave the backbone structure almost

unchanged [65], allowing us to fix the backbone during the search for the best sidechain

conformations. Second, most sidechains in high-resolution crystallographic structures

can be represented by a limited number of conformers that comply with stereochemical

and energetic constraints [66]. This observation motivated Ponder and Richards to

develop the first library of sidechain rotamers for the 17 types of residues with dihedral

angle degrees of freedom in their sidechains, based on 10 high-resolution protein

structures determined by X-ray crystallography [67]. Subsequently, a number of

additional libraries have been derived [68 – 72].

Rotamers on a fixed backbone are often used when all the sidechains need to be

modeled on a given backbone. This approach reduces the combinatorial explosion

associated with a full conformational search of all the sidechains, and is applied by some

comparative modeling [27] and protein design approaches [73]. However, ,15% of the sidechains can not be represented well by these libraries [74]. In addition, it has been

shown that the accuracy of sidechain modeling on a fixed backbone decreases rapidly

when the backbone errors are larger than 0.5 Å [75].

Earlier methods for sidechain modeling often put less emphasis on the energy or

scoring function. The function was usually greatly simplified, and consisted of the

empirical rotamer preferences and simple repulsion terms for non-bonded contacts [69].

Nevertheless, these approaches have been justified by their performance. For example, a

method based on a rotamer library compared favorably with that based on a molecular

mechanics force field, and new methods continue to be based on the rotamer library

approach [72,76,77]. The various optimization approaches include a Monte Carlo

simulation, simulated annealing, a combination of Monte Carlo and simulated annealing,

Comparative Protein Structure Modeling and its Applications to Drug Discovery 263

the dead-end elimination theorem, genetic algorithms, neural network with simulated

annealing, mean field optimization, and combinatorial searches [69,78 – 86]. Several

recent papers focused on the testing of more sophisticated potential functions for

conformational search [86,87] and development of new scoring functions for side chain

modeling [77,88], reporting higher accuracy than earlier studies.

6. COMPARATIVE MODELING BY MODELLER

MODELLER is a computer program for comparative protein structure modeling [30,34].

In the simplest case, the input is an alignment of a sequence to be modeled with the

template structures, the atomic coordinates of the templates, and a short script file.

MODELLER then automatically calculates a model containing all non-hydrogen atoms,

without any user intervention and within minutes on a Pentium processor.

MODELLER implements comparative protein structure modeling by satisfaction of

spatial restraints [30]. The spatial restraints include (i) homology-derived restraints on

the distances and dihedral angles in the target sequence, extracted from its alignment with

the template structures [30], (ii) stereochemical restraints such as bond length and bond

angle preferences, obtained from the CHARMM-22 molecular mechanics force-field

[89], (iii) statistical preferences for dihedral angles and non-bonded inter-atomic

distances, obtained from a representative set of known protein structures [90], and (iv)

optional manually curated restraints, such as those from NMR spectroscopy, rules of

secondary structure packing, cross-linking experiments, fluorescence spectroscopy,

image reconstruction from electron microscopy, site-directed mutagenesis, and intuition.

The spatial restraints, expressed as probability density functions, are combined into an

objective function that is optimized by a combination of conjugate gradients and

molecular dynamics with simulated annealing. This model building procedure is similar

to structure determination by NMR spectroscopy.

Apart from model building, MODELLER can perform additional auxiliary tasks,

including alignment of two protein sequences or their profiles, multiple alignment of

protein sequences and/or structures, calculation of phylogenetic trees, and de novo

modeling of loops in protein structures [34].

7. PHYSICS-BASED APPROACHES TO COMPARATIVE MODEL CONSTRUCTION AND REFINEMENT

In principle, an accurate and efficient method for estimating the free energy of a given

protein conformation could substantially improve the accuracy of comparative models.

That is, an accurate energy function combined with efficient sampling could be used to

refine initially constructed comparative models and thus decrease their RMS error,

assuming that the native state represents the lowest free energy state. We suggest that

accurate energy-based scoring may be particularly important for accurately reproducing

the fine details (e.g., specific hydrogen bonding interactions) of protein active sites, which

in turn will be critical for success of virtual screening using comparative models.

Two key challenges confronting this approach are (i) efficient but accurate methods for

M. Jacobson and A. Sali264

treating solvent (many methods have entirely ignored solvent or used low-accuracy

but efficient distance-dependent dielectric representations); and (ii) the estimation of

entropic contributions to free energy differences among states, which generally requires

extensive sampling.

As discussed above, loop and sidechain prediction algorithms rely on scoring functions

to guide the sampling, aiming to identify favorable conformations and reject unfavorable

ones. These scoring functions typically do not explicitly attempt to estimate the free energy

of a given conformation. Rather, most scoring functions have been based on statistical

analyses of native protein structures encoded in the so-called potentials-of-mean-force,

heuristic functional forms, or highly simplified energetic models (e.g., only van der Waals

energy terms). Electrostatic interactions and, especially, the effect of solvent are

infrequently used in such algorithms, largely due to their computational expense; a large

number of conformations of a protein model must typically be scored during comparative

model construction, and thus the scoring function must be rapid to compute.

Nonetheless, a number of groups have made efforts to use more sophisticated energetic

scoring functions for comparative model construction and refinement. Recently, the

convergence of increased computing power as well as accurate and efficient implicit

solvent models (i.e., Poisson – Boltzmann and Generalized Born models) has bolstered

these efforts. In the realm of loop modeling, there have been several groups reporting

recent studies [62,91 – 93]; all four of these studies employed all-atom force fields and

Generalized Born solvent models for scoring. An alternative recent approach is that of

Hornak and Simmerling, which uses molecular dynamics methods [94]. Xiang and Honig

developed a function that attempts to mimic the entropic contribution to free energy

without rigorous sampling [72]. These studies have focused on reproducing loop

conformations in native protein structures.

The problem of refining comparative models to improve accuracy is more challenging,

and early attempts to use molecular dynamics or energy minimization in this context had

mixed success, frequently increasing the RMS error of the models [95,96]. Nonetheless,

physics-based energy functions, although certainly not perfect, have shown impressive

abilities to distinguish native from non-native protein structures [61,97], suggesting that

the model accuracy may be currently limited by incomplete sampling rather than the

accuracy of the scoring function. A few recent reports about successful refinement of

comparative models by restrained molecular dynamics [98,99] support this viewpoint.

A complete software package for physics-based comparative model construction and

refinement has been developed [146]. The energy function employed is based on the OPLS

all-atom force field and Surface Generalized Born solvent model with a non-polar

estimator [100 – 103]. The sampling algorithms, which all use this energy function, include

side chain and loop and helix prediction; the latter capability is novel and addresses the fact

that corresponding helices in homologous proteins frequently adopt differing confor-

mations, especially at relatively low sequence identity [64,87,104]. The loop prediction

algorithm is, to our knowledge, the most accurate yet reported in the literature when tested

by reconstructing hundreds of loops in native structures. Of particular relevance to virtual

ligand screening against comparative models (as discussed below), the comparative model

construction algorithm permits inclusion of co-factors and ligands as the model is built,

which can help to improve accuracy of binding site conformations.

Comparative Protein Structure Modeling and its Applications to Drug Discovery 265

8. ACCURACY OF COMPARATIVE MODELS

The accuracy of the predicted model determines the information that can be extracted from

it. Thus, estimating the accuracy of 3D protein models in the absence of the known

structures is essential for interpreting them. The model can be evaluated as a whole as well

as in the individual regions. There are many model evaluation programs and servers [147].

The accuracy of comparative modeling is related to the percentage sequence identity on

which the model is based, correlating with the relationship between the structural and

sequence similarities of two proteins [6,96,105]. High accuracy comparative models are

generally based on more than 50% sequence identity to their templates. They tend to

have approximately 1 Å RMS error for the main-chain atoms, which is comparable to

the accuracy of a medium resolution NMR structure or a low-resolution X-ray structure.

The errors are mostly mistakes in side-chain packing, small shifts or distortions of the

core main-chain regions, and occasionally larger errors in loops. Medium accuracy

comparative models are based on 30 – 50% sequence identity. They tend to have

approximately 90% of the main-chain modeled with 1.5 Å RMS error. Errors in side-chain

packing, core distortion, and loop modeling errors are more frequent, and there are

occasional alignment mistakes [105]. Finally, low accuracy comparative models are based

on less than 30% sequence identity. Alignment errors increase rapidly below 30%

sequence identity and become the most significant source of errors in comparative models.

In addition, when a model is based on an almost insignificant alignment to a known

structure, it may also have an entirely incorrect fold. Accuracies of the best model building

methods are relatively similar when used optimally [96,106]. Other factors such as

template selection and alignment accuracy usually have a larger impact on the model

accuracy, especially for models based on less than 40% sequence identity to the templates.

9. MODELING ON A GENOMIC SCALE

Threading and comparative modeling methods have been applied on a genomic scale [105,

107,108]. Domains in approximately one half of all 1,300,000 known protein sequences

were modeled with MODPIPE [105,107] and MODELLER [30], and deposited into a

comprehensive database of comparative models, MODBASE [9]. The web interface to the

database allows flexible querying for fold assignments, sequence-structure alignments,

models, and model assessments. An integrated sequence/structure viewer, Chimera [110],

allows inspection and analysis of the query results. MODBASE is inter-linked with other

applications and databases such that structures and other types of information can be easily

used for functional annotation. For example, MODBASE contains binding site predictions

for small ligands and a set of predicted interactions between pairs of modeled sequences

from the same genome. Other resources associated with MODBASE include a

comprehensive database of multiple protein structure alignments (DBALI) as well as

web servers for automated comparative modeling with MODPIPE (MODWEB), modeling

of loops in protein structures (MODLOOP), and predicting functional consequences of

single nucleotide polymorphisms (SNPWEB) [109,111,113].

While the current number of modeled proteins may look impressive given the

early stage of structural genomics, usually only one domain per protein is modeled

M. Jacobson and A. Sali266

(on the average, proteins have slightly more than two domains) and two thirds of the

models are based on less than 30% sequence identity to the closest template.

10. APPLICATIONS OF COMPARATIVE MODELING TO DRUG DISCOVERY

There is a wide range of applications of protein structure models [5]. For example, high

and medium accuracy comparative models frequently are helpful in refining functional

predictions that have been based on a sequence match alone because ligand binding is

more directly determined by the structure of the binding site than by its sequence. It is

often possible to correctly predict features of the target protein that do not occur in the

template structure. The size of a ligand may be predicted from the volume of the binding

site cleft, and the location of a binding site for a charged ligand can be predicted from a

cluster of charged residues on the protein. Fortunately, errors in the functionally

important regions in comparative models are many times relatively low because the

functional regions, such as active sites, tend to be more conserved in evolution than the

rest of the fold. Even low accuracy comparative models may be useful, for example, for

assigning the fold of a protein. Fold assignment can be very helpful in drug discovery,

because it can shortcut the search for leads by pointing to compounds that have been

previously developed for other members of the same family [114,115].

10.1. Comparative models vs experimental structures in virtual screening

The remainder of this review focuses on the use of comparative models for ligand

docking [116 – 118]. It is widely accepted that docking to comparative models is more

challenging and less successful than docking to crystallographic structures. However,

surprisingly little work has been done to quantify the accuracy of docking to comparative

models, to determine in detail why the results are inferior to those obtained with crystal

structures, and to improve methods for docking to comparative models. Efforts along

these lines, directed at the important kinase and GPCR drug targets, are described in

separate sections below.

We begin our discussion with a study that compared the success of docking against ten

enzymes, each in three different conformations: holo (ligand bound), apo, and homology

modeled [119]. Each enzyme had multiple known inhibitors in the MDDR database, a

library of drug-like molecules where each molecule has been annotated by the receptor to

which it binds. Success of the docking, carried out with the this group’s version of

DOCK, was assessed by enrichment: the ability to distinguish known inhibitors from a

large set of ‘decoys’ (,100,000 of them in this case), relative to random selection. As might be expected, the holo structures were the best at selecting the known ligands from

among the MDDR decoys based on docking score. Unexpectedly, the comparative

models, all taken from the MODBASE website without any special preparation, often

ranked known ligands among the top-scoring database molecules. In four targets, this

enrichment was better than 20-fold over random for top-scoring molecules [119]. In one

Comparative Protein Structure Modeling and its Applications to Drug Discovery 267

case, purine nucleoside phosphorylase, the modeled structure actually performed better

than the holo structure. For the comparative model, 25% of the known ligands were found

in the top 1.2% of the ranked database, whereas for the holo conformation 2.8% of the

ranked list had to be searched before 25% of the ligands were found. Thus, whereas X-ray

crystallographic structures remain the first choice in docking, in many cases comparative

models seem good enough to highly rank known ligands from among a very large list of

possible alternatives.

10.2. Use of comparative models to obtain novel drug leads

Comparative models have been used in conjunction with virtual screening to successfully

identify novel inhibitors over the past few years. We briefly review a few of these

‘success stories’ to highlight the potential of this approach. The role of comparative

modeling in the design of new kinase inhibitors is discussed separately, below.

A series of papers demonstrate the successful use comparative protein structure

models to aid rational drug design against parasites. Comparative models were used

for computational docking studies that identified low micromolar nonpeptidic

inhibitors of proteases in malaria and the schistosome [53]. Subsequently, further

work applied similar methods to develop nanomolar inhibitors of falcipain that are

active against chloroquine-resistant strains of malaria [120]. Additionally, comparative

models were used to predict new nonpeptide inhibitors of cathepsin L-like cysteine

proteases in L. major[121]. Sixty-nine compounds were selected by DOCK 3.5 as

strong binders to a comparative model of protein cpB, and of these, 21 had

experimental IC50 values below 100 mM. Finally, comparative models were used to

rationalize ligand binding affinities of cysteine proteases in E. histolytica [122].

Specifically, this work provided an explanation for why proteins ACP1 and ACP2 had

substrate specificity similar to that of cathepsin B, although their overall structure is

more similar to that of cathepsin D.

Fifteen new inhibitors of matriptase have been discovered by docking against its

comparative model [123]. The comparative model employed thrombin as the template,

sharing only 34% sequence identity with the target sequence. Moreover, some residues in

the binding site are significantly different, including a trio of Asp residues in thrombin

that are modified to Tyr and Trp in matriptase. Thrombin was chosen as the template, in

part, because it prefers substrates with positively charged residues at the P1 position, as

does matriptase. The comparative model was constructed using MODELLER and refined

with MD simulations in CHARMM. The NCI database was used for virtual screening that

targeted the S1 site with the DOCK program. The 2000 best-scoring compounds were

manually inspected to identify positively charged ligands (the S1 site is negatively

charged), and 69 compounds were experimentally screened for inhibition, identifying 15

inhibitors. One of them, hexamidine, was used as a lead to identify additional compounds

selective for matriptase relative to thrombin. Another group has also used similar

methods to discover seven new, low-micromolar inhibitors of Bcl-2, using a comparative

model based on the NMR solution structure of Bcl-XL [124].

A novel inhibitor of a retinoic acid receptor by virtual screening using a comparative

model [125]. In this case, the target (RAR-a) and template (RAR-g) are very closely

M. Jacobson and A. Sali268

related; only three residues in the binding site are not conserved. The ICM program was

used for virtual screening of ligands from the Available Chemicals Directory, using a

hierarchical approach. As with other studies, 300 good-scoring compounds from

the virtual screening were manually inspected to choose the final 30 for testing. Two

novel agonists were identified, with 50 nanomolar activity.

Novel inhibitors of DHFR in T. cruzi (the parasite that causes Chagas’ disease) were

discovered by docking into a comparative model based on ,50% sequence identity to DHFR in L. major, a related parasite [126]. The virtual screening procedure used DOCK

for rigid docking of over 50,000 selected compounds from the Cambridge Structural

Database. Visual inspection of the top 100 hits was used to select 36 compounds for

experimental testing. This work identified several novel scaffolds with micromolar IC50 values. The authors report attempting to use virtual screening results to identify

compounds with greater affinity for T. cruzi DHFR than human DHFR, but it is not clear

how successful they were.

10.3. Comparative models of kinases in virtual screening

Protein kinases have been intensely investigated as drug targets. The large number of

kinases in the human genome (,500) is both an asset and a liability for drug development efforts. The catalytic domain of kinases (i.e., the domain that binds ATP and

effects phosphorylation of substrates) is well conserved structurally, making it possible to

construct comparative models for this domain for many of the kinases in the human

genome. On the other hand, the high level of sequence conservation around the ATP

binding site-the site most frequently targeted for inhibitor development-raises concerns

about selectivity. Drugs that inadvertently inhibit kinases unrelated to the disease state

one intends to treat may lead to side effects (although activity against several kinases may

be desirable for treating some diseases). Proteins other than kinases can also contain ATP

binding sites, of course, although it is not clear whether selectivity across protein families

is a problem in practice. Reliable comparative models of kinases can in principle help to

address issues of selectivity and aid drug development against the many kinases lacking

experimental structures. Work along these lines remains at an early stage. Here we

highlight five studies in which kinase comparative models were used in virtual screening

applications; two of these studies developed new kinase inhibitors, while the remaining

three focused on methodological issues.

A selective inhibitor of human CK2 was discovered by virtual screening [144].

Because no experimental structure of the human protein was available at the time, a

comparative model of this kinase based on the structure of CK2 from Zea mays, using the

WhatIf program [127]. The level of sequence identity between the two proteins is very

high: 72% overall, and 82% in the ATP binding site. The crystal structure of human CK2

was subsequently solved, and the Ca RMS error of the model was determined to be only

0.92 Å overall, and 0.64 Å in the binding site. DOCK was used to screen a corporate

database of 400,000 compounds against the comparative model. As with many other

studies, the hit list was subjected to extensive filtering based on a variety of factors

including visual inspection. A dozen compounds were tested, yielding several inhibitors

with activity at 10 mM, and one with an IC50 of 80 nM – the most potent inhibitor of CK2

Comparative Protein Structure Modeling and its Applications to Drug Discovery 269

yet reported. Fortuitously, this inhibitor has excellent selectivity, showing much weaker

inhibition of other kinases.

In another successful use of kinase comparative models for inhibitor development, was

in the virtual screening of a model of CDK4, a member of the important cyclic-dependent

kinase class of drug targets (CDKs play a central role in regulation of the cell cycle)

[145]. The comparative model was built from the crystal structure of CDK2 in its

activated (phosphorylated) form (45% sequence identity). The de novo design program

LEGEND was then used to predict inhibitors, which were subsequently filtered based on

commercial availability. Of 382 tested compounds, 18 had IC50 values better than

500 mM. After several rounds of optimization, one inhibitor had an IC50 of 42 nM. The

predicted binding mode of this inhibitor was subsequently supported by a crystal

structure of the inhibitor bound to CDK2.

Two recent papers have evaluated the enrichment of known kinase inhibitors (relative

to randomly selected drug-like molecules) obtained by virtual screening against kinase

comparative models. Construction of a total of 17 comparative models of six kinases,

using templates with sequence identity ranging from 30 to 77% was used in the first study

[128]. Four of these kinases had crystal structures available for comparison. Known

inhibitors for the kinases (ranging from 46 for VEGFR1 to 387 for EGFr) were combined

with 32,000 random compounds for the virtual screening, which used the LibDock

program [129]. In five of six cases, the known inhibitors were found to be enriched by

factors of 4 – 5 in the top 5% of the docking hits. These results are worse than typical

results with kinase crystal structures (although for SRC, the comparative model

performed much better than the crystal structure, for reasons that are not entirely clear).

This study also compared poses of the docked ligands with crystallographically

determined positions, when available, and suggested structural reasons for the difference

in enrichment between crystal structures and comparative models.

A further study compared enrichment for models of CDK2 built from several different

templates with results for the crystal structure [130]. Comparative models were

constructed with MOE, and docking was performed with DOCK. This study concluded

that comparative models built from templates with .50% sequence identity provide

useful proxies of crystal structures, generating enrichment factors of ,5, similar to the results with the crystal structure. Results for virtual screening against comparative

models of factor VIIa were also presented and are similar to those for CDK2.

In a study investigating the differential binding of the drugs Gleevec, purvalanol A,

and hymenialdisine to their intended kinase targets and many other kinases, most receptor

structures were constructed by comparative modeling [131]. This study represents an

early effort toward using comparative models to help assess possible side effects of drugs

caused by unintended binding to homologs of the intended target protein.

10.4. GPCR comparative models for drug development

Comparative modeling takes on special importance for the G-protein coupled receptors

(GPCRs). This class of membrane proteins has played an enormously important role in

drug development; approximately 40% of all drugs target a member of this class.

However, they are very difficult to crystallize, and most of the comparative modeling

M. Jacobson and A. Sali270

effort has been based on the atomic-resolution structure of bovine rhodopsin.

Comparative modeling methods for GPCRs capable of accuracy sufficient for

structure-based drug design would have an enormous impact on drug discovery. Several

serious challenges confront this goal. Many targets of interest for drug development share

rather low sequence identity (,30%) with rhodopsin. Although the transmembrane

helices can frequently be aligned with reasonable certainty (aided by certain highly

conserved residues), the extracellular loops are much more divergent in sequence.

Finally, model building algorithms typically do not take the membrane environment into

account explicitly, although it is not clear whether this approximation is a significant

factor for reconstructing the ligand binding sites. Despite these challenges, encouraging

early results have been reported for docking into comparative models of GPCRs. In fact,

enrichment of known ligands in virtual screening has become an important means of

validating GPCR comparative models.

A rather extensive test of docking methods with rhodopsin-based comparative

models, using three different docking programs and seven different scoring functions

has recently been reported [132]. The comparative modeling protocol is intricate, and

includes manual adjustment of sequence alignments, specialized loop search methods

for the second extracellular loop, and energy minimization of the model using AMBER.

Models were separately prepared for agonists and antagonists; that is, known ligands

were manually ‘docked’ into the binding site in reasonable poses, and the protein

structure was again minimized. Several different docking protocols were attempted, and

the authors report that the use of a ‘consensus’ scoring procedure involving three

scoring functions provides the best results. Under optimal conditions, the rather

elaborate procedures produced encouraging results in cross-docking experiments

involving known ligands of the b2 and D3 receptors, and in enrichment tests using

990 randomly chosen compounds from the ACD.

The ICM flexible docking algorithm could successfully reproduce the conformation of

the retinal ligand in rhodopsin and rank this ligand in the top 2% relative to a database of

decoy ligands [133]. Encouragingly, the method appears to be robust with respect to

entirely deleting the extracellular loop regions (which are difficult to model) and errors in

the binding site side chain conformations. GPCR models generated by the MembStruk

[134] and PREDICT [135,136] methods, which do not explicitly rely on the rhodopsin

structure, also appear to be useful for virtual screening purposes.

10.5. Other uses of comparative models in drug development

Clearly, as the previous discussion shows, comparative models are finding increasing

application in virtual ligand screening. However, comparative models have been put to

use in many other ways in the context of drug development, generally to obtain

qualitative insights to guide inhibitor development. In contrast to the literature on

docking to comparative models, the number of papers in this category is quite large, and

we aim to provide only a flavor of the many ways in which comparative models have

been used.

One common use of comparative models is to rationalize trends in binding affinities

among known inhibitors or structure-activity relationships, with an aim of guiding future

Comparative Protein Structure Modeling and its Applications to Drug Discovery 271

lead discovery and optimization efforts. In these studies, docking algorithms are used

primarily to generate possible binding modes of known inhibitors in comparative models,

providing a structural context for understanding existing data on these inhibitors.

Examples include recent studies of the cysteine protease falcipain-3 in malaria,

thymidine phosphorylase, thymidine kinase, and androgen receptor [137 – 140]. Similar

studies have also combined comparative modeling and docking to help guide

the improvement of the pharmacokinetic properties of inhibitors, rationalize the role of

mutations in the development of drug resistance, and understand selectivity of drugs

against homologous proteins [141 – 143].

10.6. Future directions

Although recent reports of successful virtual screening against comparative models are

encouraging, such efforts are not yet a routine part of rational drug design efforts. Even

the successful efforts appear to rely strongly on visual inspection of the docking results,

and it is clear that much work remains to improve the robustness and accuracy of docking

against comparative models. Although a number of studies have compared docking

against comparative models with docking against crystal structures of the same proteins,

little work has been done to compare the accuracy achievable by different approaches to

comparative modeling, or to identify the specific structural reasons why comparative

models produce less accurate virtual screening results. Among the many issues that

deserve consideration are the following:

† The inclusion of co-factors and bound waters in protein receptors is often critical for

success of virtual screening; however, co-factors are not routinely included in

comparative models.

† The accuracy of comparative models is frequently judged by the Ca RMS error, or

similar measures of backbone accuracy. For virtual screening, however, the precise

positioning of side chains in the binding site is likely to be critical; measures of

accuracy for binding sites are needed to help evaluate the suitability of comparative

modeling algorithms for constructing models for docking.

† Knowledge of known inhibitors, either for the target protein or the template, can be

used to evaluate and improve virtual screening against comparative models. For

example, comparative models constructed from ‘holo’ template structures implicitly

preserve some information about the ligand-bound receptor conformation.

† Improvement in the accuracy of models produced by comparative modeling will

require methods that finely sample protein conformational space using a free energy or

scoring function that has sufficient accuracy to distinguish the native structure from

the non-native conformations. Despite many years of development of molecular

simulation methods, attempts to refine models that are already relatively close to the

native structure have met with relatively little success. This failure is likely to be due

to inaccuracies in the potential functions used in the simulations, particularly in the

treatment of electrostatics and solvation effects. Improvements in sampling strategies

may also be necessary. Combination of physical chemistry with the vast amount of

M. Jacobson and A. Sali272

information in known protein structures may provide a route to development of

improved potential functions.

11. CONCLUSIONS

High-resolution crystallographic structures becoming available for all human and

pathogen proteins and complexes is a very unlikely scenario. The only practical manner

of exploring ligand – protein interactions for most systems is to use comparative protein

structure models. It is now clear that comparative models, based on as little as 30%

sequence identity to known template structures, can be useful. New methods that improve

and combine existing modeling and docking techniques to further advance the utility of

comparative models will no doubt be developed in the future.

REFERENCES

[1] S. E. Brenner, Nat. Struct. Biol., 2000, 7(Suppl.), 967. [2] B. Boeckmann, A. Bairoch, R. Apweiler, M. C. Blatter, A. Estreicher, E. Gasteiger, M. J.

Martin, K. Michoud, C. O’Donovan, I. Phan, et al., Nuclic Acids Res., 2003, 31, 365. [3] D. A. Benson, I. Karsch-Mizrachi, D. J. Lipman, J. Ostell, B. A. Rapp and L. Wheeler,

Nucleic Acids Res., 2002, 30, 17. [4] H. M. Berman, T. Battistuz, T. N. Bhat, W. F. Bluhm, P. E. Bourne, K. Burkhardt, Z. Feng,

G. L. Gilliland, L. Iype, S. Jain, et al., Acta Crystallogr. D., 2002, 58, 899. [5] D. Baker and A. Sali, Science, 2001, 294, 93. [6] M. A. Marti-Renom, A. Stuart, A. Fiser, R. Sanchez, F. Melo and A. Sali, Annu. Rev.

Biophys. Biomol. Struct., 2000, 29, 291. [7] R. Bonneau and D. Baker, Annu. Rev. Biophys. Biomol. Struct., 2001, 30, 173. [8] D. Baker, Nature, 2000, 405, 39. [9] U. Pieper, N. Eswar, H. Braberg, M. S. Madhusudhan, F. P. Davis, A. C. Stuart, N. Mirkovic,

A. Rossi, M. A. Marti-Renom, A. Fiser, et al., Nucleic Acids Res., 2004, 32, D217, Database issue, http://www.salilab.org/modbase/ and http://www.nysqxrc.org/.

[10] D. Fischer, A. Elofsson, L. Rychlewski, F. Pazos, A. Valencia, B. Rost, A. R. Ortiz and R. L. Dunbrack, Jr, Proteins, 2001, 45 (Suppl. 5), 171.

[11] P. E. Bourne, Methods Biochem. Anal., 2003, 44, 501. [12] J. M. Bujnicki, A. Elofsson, D. Fischer and L. Rychlewski, Prot. Sci., 2001, 10, 352. [13] V. A. Eyrich, M. A. Marti-Renom, D. Przybylski, M. S. Madhusudhan, A. Fiser, F. Pazos,

A. Valencia, A. Sali and B. Rost, Bioinformatics, 2001, 17, 1242. [14] I.-Y. Y. Koh, V. A. Eyrich, M. A. Marti-Renom, D. Przybylski, M. S. Madhusudhan,

E. Narayanan, O. Graca, F. Pazos, A. Valencia, A. Sali, et al., Nucleic Acids Res., 2003. [15] A. Sali, Nat. Struct. Biol., 1998, 5, 1029. [16] S. K. Burley, S. C. Almo, J. B. Bonanno, M. Capel, M. R. Chance, T. Gaasterland, D. Lin,

A. Sali, F. W. Studier and S. Swaminathan, Nat. Genet., 1999, 23, 151. [17] D. Vitkup, E. Melamud, J. Moult and C. Sander, Nat. Struct. Biol., 2001, 8, 559. [18] S. F. Altschul, W. Gish, W. Miller, E. W. Myers and D. J. Lipman, J. Mol. Biol., 1990, 215,

403. [19] W. R. Pearson, Methods Enzymol., 1990, 183, 63. [20] S. F. Altschul, T. L. Madden, A. A. Schaffer, J. Zhang, Z. Zhang, W. Miller and D. J. Lipman,

Nucleic Acids Res., 1997, 25, 3389. [21] M. A. Marti-Renom, M. S. Madhusudhan and A. Sali, Prot. Sci., 2004, 13, 1071. [22] R. Karchin, M. Cline, Y. Mandel-Gutfreund and M. Karplus, Proteins, 2003, 51, 504.

Comparative Protein Structure Modeling and its Applications to Drug Discovery 273

[23] A. E. Torda, Curr. Opin. Struct. Biol., 1997, 7, 200. [24] B. John and A. Sali, Nucleic Acids Res., 2003, 31, 3982 – 3992. [25] W. J. Browne, A. C. T. North, D. C. Phillips, K. Brew, T. C. Vanaman and R. C. Hill, J. Mol.

Biol., 1969, 42, 65. [26] J. Greer, J. Mol. Biol., 1981, 153, 1027. [27] T. L. Blundell, B. L. Sibanda, M. J. Sternberg and J. M. Thornton, Nature, 1987, 326, 347. [28] M. Levitt, J. Mol. Biol., 1992, 226, 507. [29] T. F. Havel and M. E. Snow, J. Mol. Biol., 1991, 217, 1. [30] A. Sali and T. L. Blundell, J. Mol. Biol., 1993, 234, 779. [31] A. Kolinski, M. R. Betancourt, D. Kihara, P. Rotkiewicz and J. Skolnick, Proteins, 2001, 44,

133. [32] W. Kabsch and C. Sander, Proc. Natl. Acad. Sci. USA, 1984, 81, 1075. [33] M. Mezei, Prot. Eng., 1998, 11, 411. [34] A. Fiser, R. K. Do and A. Sali, Prot. Sci., 2000, 9, 1753, http://www.salilab.org/modeller. [35] M. P. Jacobson, D. L. Pincus, C. S. Rapp, T. J. Day, B. Honig, D. E. Shaw and R. A. Friesner,

Proteins, 2004, 55, 351. [36] T. A. Jones and S. Thirup, EMBO J., 1986, 5, 819. [37] C. Chothia and M. Lesk, J. Mol. Biol., 1987, 196, 901. [38] J. Moult and M. N. James, Proteins, 1986, 1, 146. [39] R. E. Bruccoleri and M. Karplus, Biopolymers, 1987, 26, 137. [40] P. S. Shenkin, D. L. Yarmush, R. M. Fine, H. J. Wang and C. Levinthal, Biopolymers, 1987,

26, 2053. [41] H. W. van Vlijmen and M. Karplus, J. Mol. Biol., 1997, 267, 975. [42] C. M. Deane and T. L. Blundell, Prot. Sci., 2001, 10, 599. [43] B. L. Sibanda, T. L. Blundell and J. M. Thornton, J. Mol. Biol., 1989, 206, 759. [44] C. Chothia, A. M. Lesk, A. Tramontano, M. Levitt, S. J. Smith-Gill, G. Air, S. Sheriff, E. A.

Padlan, D. Davies, W. R. Tulip, et al., Nature, 1989, 342, 877. [45] C. S. Ring, D. G. Kneller, R. Langridge and F. E. Cohen, J. Mol. Biol., 1992, 224, 685. [46] B. Oliva, P. A. Bates, E. Querol, F. X. Aviles and M. J. Sternberg, J. Mol. Biol., 1997, 266,

814. [47] S. D. Rufino, L. E. Donate, L. H. Canard and T. L. Blundell, J. Mol. Biol., 1997, 267, 352. [48] K. Fidelis, P. S. Stern, D. Bacon and J. Moult, Prot. Eng., 1994, 7, 953. [49] U. Lessel and D. Schomburg, Prot. Eng., 1994, 7, 1175 – 1187. [50] M. Claessens, E. Van Cutsem, I. Lasters and S. Wodak, Prot. Eng., 1989, 2, 335. [51] R. M. Fine, H. Wang, P. S. Shenkin, D. L. Yarmush and C. Levinthal, Proteins, 1986, 1, 342. [52] R. E. Bruccoleri and M. Karplus, Biopolymers, 1990, 29, 1847. [53] C. S. Ring, E. Sun, J. H. McKerrow, G. K. Lee, P. J. Rosenthal, I. D. Kuntz and F. E. Cohen,

Proc. Natl Acad. Sci. USA, 1993, 90, 3583. [54] J. Higo, V. Collura and J. Garnier, Biopolymers, 1992, 32, 33. [55] V. Collura, J. Higo and J. Garnier, Prot. Sci., 1993, 2, 1502. [56] R. Abagyan and M. Totrov, J. Mol. Biol., 1994, 235, 983. [57] Q. Zheng, R. Rosenfeld, S. Vajda and C. DeLisi, Prot. Sci., 1993, 2, 1242. [58] P. Koehl and M. Delarue, Nat. Struct. Biol., 1995, 2, 163. [59] R. Samudrala and J. Moult, J. Mol. Biol., 1998, 279, 287. [60] Z. Xiang, C. S. Soto and B. Honig, Proc. Natl Acad. Sci. USA, 2002, 99, 7432 – 7437. [61] A. K. Felts, E. Gallicchio, A. Wallqvist and R. M. Levy, Proteins, 2002, 48, 404. [62] P. I. de Bakker, M. A. DePristo, D. F. Burke and T. L. Blundell, Proteins, 2003, 51, 21. [63] M. A. DePristo, P. I. de Bakker, S. C. Lovell and T. L. Blundell, Proteins, 2003, 51, 41. [64] M. P. Jacobson, D. L. Pincus, C. S. Rapp, T. J. F. Day, B. Honig, D. E. Shaw and R. A.

Friesner, Proteins, 2004, 55, 351. [65] C. Chothia, A. M. Lesk, M. Levitt, A. G. Amit, R. A. Mariuzza, S. E. Phillips and R. J.

Poljak, Science, 1986, 233, 755. [66] J. Janin and C. Chothia, Biochemistry, 1978, 17, 2943. [67] J. W. Ponder and F. M. Richards, J. Mol. Biol., 1987, 193, 775. [68] P. Tuffery, C. Etchebest, S. Hazout and R. Lavery, J. Biomol. Struct. Dyn., 1991, 8, 1267.

M. Jacobson and A. Sali274

[69] R. L. Dunbrack and M. Karplus, J. Mol. Biol., 1993, 230, 543. [70] R. L. Dunbrack, Jr and F. E. Cohen, Prot. Sci., 1997, 6, 1661. [71] J. Mendes, A. M. Baptista, M. A. Carrondo and C. M. Soares, Proteins, 1999, 37, 530. [72] Z. Xiang and B. Honig, J. Mol. Biol., 2001, 311, 421 – 430. [73] J. R. Desjarlais and T. M. Handel, J. Mol. Biol., 1999, 290, 305. [74] V. De Filippis, C. Sander and G. Vriend, Prot. Eng., 1994, 7, 1203. [75] S. Y. Chung and S. Subbiah, Pac. Symp. Biocomput., 1996, 126. [76] D. Cregut, J. P. Liautard and L. Chiche, Prot. Eng., 1994, 7, 1333. [77] A. A. Canutescu, A. A. Shelenkov and R. L. Dunbrack, Jr, Prot. Sci., 2003, 12, 2001. [78] F. Eisenmenger, P. Argos and R. Abagyan, J. Mol. Biol., 1993, 231, 849. [79] G. M. Lee, A. Varma and B. O. Palsson, Biotechnol. Prog., 1991, 7, 72. [80] L. Holm and C. Sander, Proteins, 1992, 14, 213. [81] I. Lasters and J. Desmet, Prot. Eng., 1993, 6, 717. [82] L. L. Looger and H. W. Hellinga, J. Mol. Biol., 2001, 307, 429. [83] J. K. Hwang and W. F. Liao, Prot. Eng., 1995, 8, 363. [84] P. Koehl and M. Delarue, J. Mol. Biol., 1994, 239, 249. [85] M. J. Bower, F. E. Cohen and R. L. Dunbrack, Jr, J. Mol. Biol., 1997, 267, 1268. [86] R. J. Petrella, T. Lazaridis and M. Karplus, Fold. Des., 1998, 3, 353. [87] M. P. Jacobson, G. A. Kaminski, R. A. Friesner and C. S. Rapp, J. Phys. Chem. B, 2002, 106,

11673. [88] S. Liang and N. V. Grishin, Prot. Sci., 2002, 11, 322. [89] A. D. MacKerell, Jr, D. Bashford, M. Bellott, R. L. Dunbrack, Jr, J. D. Evanseck, M. J. Field,

S. Fischer, J. Gao, H. Guo, S. Ha, et al., J. Phys. Chem. B, 1998, 102, 3586. [90] A. Sali and J. P. Overington, Prot. Sci., 1994, 3, 1582 – 1596. [91] C. S. Rapp and R. A. Friesner, Proteins, 1999, 35, 173. [92] A. Fiser, M. Feig, C. L. Brooks, III and A. Sali, Acc. Chem. Res., 2002, 35, 413. [93] M. P. Jacobson, R. A. Friesner, Z. X. Xiang and B. Honig, J. Mol. Biol., 2002, 320, 597. [94] V. Hornak and C. Simmerling, Proteins, 2003, 51, 577. [95] J. Moult, Curr. Opin. Biotechnol., 1996, 7, 422. [96] P. Koehl and M. Levitt, Nat. Struct. Biol., 1999, 6, 108. [97] B. N. Dominy and C. L. Brooks, J. Comp. Chem., 2002, 23, 147. [98] H. Li, R. Tejero, D. Monleon, D. BassolinoKlimas, C. AbateShen, R. E. Bruccoleri and G. T.

Montelione, Prot. Sci., 1997, 6, 956. [99] J. A. Flohil, G. Vriend and H. J. C. Berendsen, Proteins, 2002, 48, 593.

[100] W. L. Jorgensen and J. Tiradorives, J. Am. Chem. Soc., 1988, 110, 1666. [101] G. A. Kaminski, R. A. Friesner, J. Tirado-Rives and W. L. Jorgensen, J. Phys. Chem. B,

2001, 105, 6474. [102] A. Ghosh, C. S. Rapp and R. A. Friesner, J. Phys. Chem. B, 1998, 102, 10983. [103] E. Gallicchio, L. Y. Zhang and R. M. Levy, J. Comp. Chem., 2002, 23, 517. [104] X. Li, M. P. Jacobson and R. A. Friesner, Proteins, 2004. [105] R. Sanchez and A. Sali, Proc. Natl Acad. Sci. USA, 1998, 95, 13597. [106] M. A. Marti-Renom, M. S. Madhusudhan, A. Fiser, B. Rost and A. Sali, Structure, 2002, 10, 435. [107] D. Fischer and D. Eisenberg, Proc. Natl Acad. Sci. USA, 1997, 94, 11929. [108] N. Guex, A. Diemand and M. C. Peitsch, Trends Biochem. Sci., 1999, 24, 364. [109] N. Eswar, B. John, N. Mirkovic, A. Fiser, V. A. Ilyin, U. Pieper, A. C. Stuart, M. A. Marti-

Renom, M. S. Madhusudhan, B. Yerkovich, et al., Nucleic Acids Res., 2003, 31, 3375, http:// www.salilab.org/modweb.

[110] C. C. Huang, W. R. Novak, P. C. Babbitt, A. I. Jewett, T. E. Ferrin and T. E. Klein, Pac. Symp. Biocomput., 2000, 230.

[111] M. A. Marti-Renom, V. A. Ilyin and A. Sali, Bioinformatics, 2001, 17, 746, http://www. salilab.org/dbali.

[112] A. Fiser and A. Sali, Bioinformatics, 2003, 19, 2500.Q1

[113] N. Mirkovic, M. A. Marti-Renom, A. Sali and A. N. A. Monteiro, Cancer Res., 2004, 64, 3790. [114] R. K. Gordon, K. Ginalski, W. R. Rudnicki, L. Rychlewski, M. C. Pankaskie, J. M. Bujnicki

and P. K. Chiang, Eur. J. Biochem., 2003, 270, 3507.

Comparative Protein Structure Modeling and its Applications to Drug Discovery 275

[115] M. von Grotthuss, L. S. Wyrwicz and L. Rychlewski, Cell, 2003, 113, 701. [116] A. Schafferhans and G. Klebe, J. Mol. Biol., 2001, 307, 407 – 427. [117] A. Evers, H. Gohlke and G. Klebe, J. Mol. Biol., 2003, 334, 327 – 345. [118] A. Evers and G. Klebe, Angew Chem. Int. Ed. Engl., 2004, 43, 248 – 251. [119] S. L. McGovern and B. K. Shoichet, J. Med. Chem., 2003, 46, 2895. [120] R. Li, X. Chen, B. Gong, P. M. Selzer, Z. Li, E. Davidson, G. Kurzban, R. E. Miller, E. O.

Nuzum, J. H. McKerrow, et al., Bioorg. Med. Chem. Lett., 1996, 4, 1421. [121] P. M. Selzer, X. Chen, V. J. Chan, M. Cheng, G. L. Kenyon, I. D. Kuntz, J. A. Sakanari, F. E.

Cohen and J. H. McKerrow, Exp. Parasitol., 1997, 87, 212. [122] X. Que, L. S. Brinen, P. Perkins, S. Herdman, K. Hirata, B. E. Torian, H. Rubin, J. H.

McKerrow and S. L. Reed, Mol. Biochem. Parasitol., 2002, 119, 23. [123] I. J. Enyedy, S. L. Lee, A. H. Kuo, R. B. Dickson, C. Y. Lin and S. Wang, J. Med. Chem.,

2001, 44, 1349. [124] I. J. Enyedy, Y. Ling, K. Nacro, Y. Tomita, X. Wu, Y. Cao, R. Guo, B. Li, X. Zhu, Y. Huang,

et al., J. Med. Chem., 2001, 44, 4313. [125] M. Schapira, B. M. Raaka, H. H. Samuels and R. Abagyan, BMC Struct. Biol., 2001, 1, 1. [126] F. Zuccotto, M. Zvelebil, R. Brun, S. F. Chowdhury, R. Di Lucrezia, I. Leal, L. Maes, L. M.

Ruiz-Perez, D. Gonzalez Pacanowska and I. H. Gilbert, Eur. J. Med. Chem., 2001, 36. [127] G. Vriend, J. Mol. Graph., 1990, 8, 52. [128] D. J. Diller and R. Li, J. Med. Chem., 2003, 46, 4638. [129] D. J. Diller and K. M. Merz, Proteins, 2001, 43, 113. [130] C. Oshiro, E. K. Bradley, J. Eksterowicz, E. Evensen, M. L. Lamb, J. K. Lanctot, S. Putta,

R. Stanton and P. D. Grootenhuis, J. Med. Chem., 2004, 47, 764. [131] W. M. Rockey and A. H. Elcock, Proteins, 2002, 48, 664. [132] C. Bissantz, P. Bernard, M. Hibert and D. Rognan, Proteins, 2003, 50, 5. [133] C. N. Cavasotto, A. J. Orry and R. A. Abagyan, Proteins, 2003, 51, 423. [134] N. Vaidehi, W. B. Floriano, R. Trabanino, S. E. Hall, P. Freddolino, E. J. Choi, G. Zamanakos

and W. A. Goddard, III, Proc. Natl Acad. Sci. USA, 2002, 99, 12622. [135] S. Shacham, M. Topf, N. Avisar, F. Glaser, Y. Marantz, S. Bar-Haim, S. Noiman, Z. Naor

and O. M. Becker, Med. Res. Rev., 2001, 21, 472. [136] O. M. Becker, S. Shacham, Y. Marantz and S. Noiman, Curr. Opin. Drug. Discov. Devel.,

2003, 6, 353. [137] Y. A. Sabnis, P. V. Desai, P. J. Rosenthal and M. A. Avery, Prot. Sci., 2003, 12, 501. [138] M. L. Price, W. C. Guida, T. E. Jackson, J. A. Nydick, P. L. Gladstone, J. C. Juarez, F. Donate

and R. J. Ternansky, Bioorg. Med. Chem. Lett., 2003, 13, 107. [139] L. Spadola, E. Novellino, G. Folkers and L. Scapozza, Eur. J. Med. Chem., 2003, 38, 413. [140] C. A. Marhefka, B. M. Moore, T. C. Bishop, II, L. Kirkovsky, A. Mukherjee, J. T. Dalton and

D. D. Miller, J. Med. Chem., 2001, 44, 1729. [141] M. H. Rabinowitz, R. C. Andrews, J. D. Becherer, D. M. Bickett, D. G. Bubacz, J. G.

Conway, D. J. Cowan, M. Gaul, K. Glennon, M. H. Lambert, et al., J. Med. Chem., 2001, 44, 4252.

[142] R. T. Delfino, O. A. Santos-Filho and J. D. Figueroa-Villar, Biophys. Chem., 2002, 98, 287. [143] E. A. Salter, A. Wierzbicki, G. Sperl and W. J. Thompson, Int. J. Quantum Chem., 2004, 96,

402. [144] E. Vangrevelinghe, K. Zimmermann, J. Schoepfer, R. Portmann, D. Fabbro and P. Furet,

J. Med. Chem., 2003, 46, 2656. [145] T. Honma, K. Hayashi, T. Aoyama, N. Hashimoto, T. Machida, K. Fukasawa, T. Iwama, C.

Ikeura, M. Ikuta, I. Suzuki-Takahashi, Y. Iwasawa, T. Hayama, S. Nishimura and H. Morishima, J. Med. Chem., 2001, 44, 4628.

[146] http://www.francisco.compbio.ucsf.edu/~jacobson/plop_overview.htm. [147] http://www.salilab.org/bioinformatics_resources.shtml.

M. Jacobson and A. Sali276

  • Comparative Protein Structure Modeling and its Applications to Drug Discovery
    • Introduction
    • Fold assignment and sequence-structure alignment
    • Comparative model building
    • Loop modeling
    • Sidechain modeling
    • Comparative modeling by MODELLER
    • Physics-based approaches to comparative model construction and refinement
    • Accuracy of comparative models
    • Modeling on a genomic scale
    • Applications of comparative modeling to drug discovery
      • Comparative models vs experimental structures in virtual screening
      • Use of comparative models to obtain novel drug leads
      • Comparative models of kinases in virtual screening
      • GPCR comparative models for drug development
      • Other uses of comparative models in drug development
      • Future directions
    • Conclusions
    • References