







































ARESTY RUTGERS UNDERGRADUATE RESEARCH JOURNAL, VOLUME I, ISSUE V 

 

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License. 

RELATIONSHIP BETWEEN 

BIOPHYSICAL PROPERTIES 

OF ANTIMICROBIAL PEP-

TIDES (AMPS) AND THEIR 

ASSOCIATED DRUG EFFI-

CACIES 
 

ALISHA ZHU 

 

✵  ABSTRACT 
Antibiotic resistance is a growing global 

health threat. One consequence is that patients with 

cystic fibrosis (CF) are prone to developing antibi-

otic resistant lung infections caused by multiple 

strains of bacteria, including Pseudomonas aeru-

ginosa. Due to the limited number of treatment op-

tions for patients with chronic antibiotic resistant in-

fections, there is a need for finding new antibiotics 

that allow for effective eradication of bacterial infec-

tions, such as those in the CF lung. Many antimicro-

bial peptides (AMPs) have been annotated in data-

bases and are considered as potential alternatives 

for current antibiotics. However, in many instances, 

the suitability of AMPs as drug molecules has not 

been extensively explored. Here, we propose that 

certain molecular properties of AMPs favor high an-

tibiotic efficacy. Using information from AMP data-

bases, we combined statistical analyses and ma-

chine learning techniques to identify relationships 

between various biophysical properties of AMPs 

and their drug efficacies. Analyses from classifica-

tion and regression trees (CART) and random for-

ests suggest that net charge and maximum average 

hydrophobic moment are the most important prop-

erties in determining if a peptide is useful against P. 

aeruginosa infections in CF patients. Maximum av-

erage hydrophobic residue, average alpha helix 

propensity score, hydrophobic proportion, and 

peptide length still contribute to this determination 

but to lesser degrees. Cation-pi interactions, on the 

other hand, do not appear to factor into this deci-

sion at all. Based on these properties, our current 

work is focused on designing and experimentally 

testing new peptides that may have activity against 

P. aeruginosa infections.  

 

2 INTRODUCTION 
Antibiotic resistance is a persistent issue in 

patient care that affects more than 2.8 million indi-

viduals, with an estimated occurrence of ~35,000 

deaths in the U.S. each year (CDC, 2019). Deemed 

as a serious threat by the CDC, Pseudomonas aeru-

ginosa, a type of rod-shaped, Gram-negative bacte-

ria, is a major source of infection that requires 

greater attention. Gram-negative bacterial infec-

tions tend to be difficult to treat because they have 

an extra outer protective membrane layer. Due to 

buildup of mucus and formation of bacterial bio-

films that prevent the delivery of antibiotics at the 

levels needed for bacteria eradication, patients with 

cystic fibrosis (CF) are susceptible to developing an-

tibiotic resistant lung infections caused by bacteria, 

of which the most common in adults is P. aeru-

ginosa. CF is a recessive genetic disorder that arises 

from mutations in the cystic fibrosis transmembrane 

conductance regulator (CFTR) gene. Loss of func-

tion of the CFTR decreases hydration in the airways, 

leading to decreased mucociliary clearance and 

bacterial retention (Allen et al., 2020). Conse-

quently, 80% of CF patients develop chronic infec-

tions (Moreau-Marquis et al., 2008). Thus, there is a 

need for the discovery of new antibiotic agents. An-

timicrobial peptides (AMPs) are a class of small pep-

tides that exist in nature. Because they exhibit a 

broad range of activity against various bacteria, 

AMPs hold potential as small molecule antibiotics 

(Ageitos et al., 2017).  

AMPs have various modes of action. First, 

AMPs can directly kill bacteria by disrupting the cell 

membrane. AMPs can destabilize and permeabilize 

the bacterial cell membrane through both hydro-

phobic and electrostatic interactions. In many cases, 

the positive net charge of AMPs allows them to in-

teract with the negatively charged regions of the 



ARESTY RUTGERS UNDERGRADUATE RESEARCH JOURNAL, VOLUME I, ISSUE V 

 

 

bacterial surface. These interactions often lead to 

the formation of membrane pores resulting in lysis, 

or the rupture of the cell membrane (Li et al., 2021). 

Second, AMPs can prevent the formation of bacte-

rial cell walls by blocking peptidoglycan elongation. 

Peptidoglycan is a polymer consisting of sugar 

chains interlinked with peptides. Since the bacterial 

cell wall is an essential protective barrier and rein-

forces cell shape, inhibiting cell wall biosynthesis ef-

fectively kills the bacteria (Chen et al., 2020). Third, 

AMPs can penetrate the cell membrane and act on 

intracellular targets, including RNA, DNA, and ribo-

somes. Although this mode of action does not result 

in cell lysis, it interferes with bacterial bioprocesses. 

For instance, some AMPs block translation initiation 

by binding to the P-site on ribosomes, while others 

inhibit protein synthesis by preventing RNA splicing 

post-transcription (Ageitos et al., 2017).  

Promising AMPs proposed in literature, par-

ticularly cationic antimicrobial peptides (CAMPs)—

i.e., AMPs that have an overall net positive charge—

suggest that certain characteristics of peptides favor 

activity against Gram-negative bacterial infections, 

such as those caused by P. aeruginosa (Lei et al., 

2019). Although AMPs have been documented in 

databases, in many cases, their microbiological ac-

tivity is not established. The rules governing which 

peptides are likely to be effective are not under-

stood, especially for a specific species of bacteria 

such as P. aeruginosa (Ramazi et al., 2022). Thus, we 

propose to use database analysis to explore the re-

lationship between biophysical properties of AMPs 

and their corresponding effectiveness. We define 

effectiveness based on a peptide’s minimum inhibi-

tory concentration (MIC), which is the lowest con-

centration of peptide that prevents visible bacterial 

growth. We selected the Database of Antimicrobial 

Activity and Structure of Peptides (DBAASP) for our 

studies. A set of candidate peptides designed to be 

efficacious will be identified and tested against clin-

ical strains of P. aeruginosa.  

 

3 METHODOLOGY 
DATABASE SELECTION AND FILTERING  

Numerous databases have been con-

structed to provide classified information for 

productive AMP design and research. The DBAASP 

is one of the larger AMP databases, since it contains 

information on both synthetic and natural peptides 

(Pirtskhalava et al., 2021). The inclusion of both syn-

thetic and natural peptides widens the scope of an-

timicrobial research and development by allowing 

for a more comprehensive exploration of the pep-

tide design space. At the time of use, the DBAASP 

contained information on 18,000+ peptides. How-

ever, the complete DBAASP data set file contained 

150,000+ entries; several peptides had multiple 

data entries containing results from different exper-

iments. We implemented initial conditions to ac-

quire a suitable subset. For the DBAASP, selected 

peptides must be monomers, cannot contain any D-

amino acids (as we do not wish to work with syn-

thetic amino acids), and cannot contain any uncom-

mon or variable amino acids. These conditions en-

sure that the subset consists of peptides, not pro-

teins, that are more likely to be compatible with bi-

ological systems. In addition, the peptides must 

have lengths less than or equal to 50 amino acids 

but also greater than or equal to 11 amino acids. 

Shorter lengths enable AMPs to have greater effi-

ciency and flexibility when interacting with bacterial 

cell membranes, which is why an upper limit for 

peptide length was specified. However, the AMPs 

also had to have lengths of at least 11 amino acids 

long, since this represents three turns of an alpha 

helix spanning the minimum thickness of a Gram-

negative bacterial cell membrane. Most im-

portantly, the peptide must have known activity 

against P. aeruginosa, but its MIC value cannot ex-

ceed 500 µM. A high MIC value indicates that more 

peptide is required for inhibiting growth; we de-

fined a MIC > 500 µM as ineffective for treating P. 

aeruginosa infections. If a peptide’s reported MIC 

was a lower threshold value (i.e., greater than some 

value), then we considered its MIC value to be dou-

ble its lower threshold. This was done to account for 

one more factor of two in the two-fold serial dilution 

performed in a MIC assay. If a peptide’s reported 

MIC was an upper threshold value (i.e., lower than 

some value), then we considered its MIC value to be 

equal to its upper threshold. If a peptide had numer-

ous reported MIC values, the lowest MIC was 



ARESTY RUTGERS UNDERGRADUATE RESEARCH JOURNAL, VOLUME I, ISSUE V 

 

 

selected. After these initial conditions were imple-

mented, we arrived at a subset of 4,218 unique pep-

tides.  

 

MOLECULAR PROPERTIES OF BIOPHYSICAL SIGNIFICANCE  

Based on what is known about AMPs, there 

are several molecular properties that are relatively 

simple to compute and are likely to influence AMP 

activity: 

PEPTIDE LENGTH. The peptide length is the to-

tal number of amino acid residues in the peptide se-

quence. AMPs are typically quite short; most of 

them have lengths less than 50 amino acids, but 

some can have lengths as large as 100 amino acids 

(Ageitos et al., 2017).  

NET CHARGE. A peptide’s net charge is calcu-

lated by subtracting the total number of negatively 

charged amino acid residues from the total number 

of positively charged amino acid residues.  

 

𝑁𝑒𝑡 𝐶ℎ𝑎𝑟𝑔𝑒 = 𝐿𝑦𝑠 + 𝐴𝑟𝑔 − 𝐴𝑠𝑝 − 𝐺𝑙𝑢     (1) 

 

Net charge indicates how an AMP will elec-

trostatically interact with the bacterial membrane. 

Studies have shown that decreasing the net charge 

to less than +4 rendered peptides inactive, and in-

creasing net charge improved antimicrobial activity. 

However, an increase to net charges greater than +9 

led to a dramatic rise in hemolytic activity (Jiang et 

al., 2008). This is undesirable because greater he-

molytic activity leads to red blood cell damage and 

can result in anemia, inflammation, and organ dys-

function.  

HYDROPHOBIC PROPORTION. The hydrophobic 

proportion refers to the percentage of hydrophobic 

amino acid residues within the peptide sequence. 

Hydrophobic residues facilitate interactions with the 

phospholipid fatty acyl chains of the bacterial mem-

brane (Edwards et al., 2016). AMPs typically contain 

40%-60% hydrophobic residues (Ye et al., 2019). 

We considered an amino acid to be hydrophobic if 

it exceeds 0.700 on Fauchere and Pliska’s (1983) hy-

drophobicity scale based on octanol-water partition 

coefficients, which is defined as the ratio of the pep-

tide’s concentration in the octanol phase to its con-

centration in the aqueous phase. If the coefficient 

value is greater than one, then the peptide is more 

lipophilic (dissolves in lipids); if the coefficient value 

is less than one, then the peptide is more hydro-

philic (dissolves in water).  

 

𝐻𝑝𝑟𝑜𝑝 =  
𝐶𝑦𝑠+𝑃ℎ𝑒+𝐼𝑙𝑒+𝐿𝑒𝑢+𝑀𝑒𝑡+𝑃𝑟𝑜+𝑉𝑎𝑙+𝑇𝑟𝑝+𝑇𝑦𝑟

𝑃𝑒𝑝𝑡𝑖𝑑𝑒 𝐿𝑒𝑛𝑔𝑡ℎ
     (2)  

 

MAXIMUM AVERAGE HYDROPHOBIC MOMENT. The 

hydrophobic moment is a measure of the am-

phiphilicity, or asymmetry of hydrophobicity, of a 

peptide’s structure. A large hydrophobic moment 

corresponds to a structure that is primarily hydro-

phobic on one side and primarily hydrophilic on the 

other side. Eisenberg et al. defined the hydropho-

bic dipole moment as  

 

𝜇 = √[∑ 𝐻𝑖 cos(𝑖𝛿)𝑖 ]2 + [∑ 𝐻𝑖 sin(𝑖𝛿)𝑖 ]2     (3)  

 

where δ represents the angle between the 

amino acid side chains (δ = 100 for an alpha helix), i 

represents the residue number in the ith position, 

and Hi represents the ith amino acid’s hydrophobi-

city based on an averaged consensus scale (Eisen-

berg et al., 1982, 1984). We modified this equation 

to determine the maximum average hydrophobic 

moment across an 11 amino acid window, which 

represents three turns of an alpha helix, with δ rang-

ing from 90° to 110° instead of fixing it at 100°. By 

adding these modifications, we hope to account for 

peptides that have alpha helical segments but may 

not have a complete alpha helical structure. Addi-

tionally, instead of using the averaged consensus 

scale, values from Fauchere and Pliska’s (1983) scale 

were used.  

 

< 𝜇𝑚𝑎𝑥

>  =
√[∑ 𝐻𝑖 cos(𝑖𝛿)𝑖 ]2 + [∑ 𝐻𝑖 sin(𝑖𝛿)𝑖 ]2

11
     (4) 

 

MAXIMUM AVERAGE HYDROPHOBIC RESIDUE. The 

hydrophobic residue calculation provides a meas-

ure of the average hydrophobicity across an 11 

amino acid window. As in the maximum average hy-

drophobic moment calculation, Hi represents the ith 

amino acid’s hydrophobicity based on Fauchere 

and Pliska’s scale (1983).  



ARESTY RUTGERS UNDERGRADUATE RESEARCH JOURNAL, VOLUME I, ISSUE V 

 

 

 

< 𝐻𝑚𝑎𝑥 >  =
∑ 𝐻𝑖𝑖

11
     (5) 

 

AVERAGE ALPHA HELIX PROPENSITY SCORE. Alpha-

helices are the most common secondary structure 

seen in AMPs. Studies have demonstrated that in-

creased alpha-helical content correlates to greater 

antimicrobial activity (Li et al., 2021). The alpha helix 

propensity score describes how likely a peptide will 

form an alpha helix. Pace and Scholtz (1998) created 

a helix propensity scale based on experimental 

studies of proteins and peptides. The propensity 

score is given in terms of the amount of energy re-

quired for a given amino acid to fold into an alpha 

helix, with a higher propensity score corresponding 

to a smaller probability that the peptide will form an 

alpha helical shape. We modified Porto et al.’s 

(2018) alpha helix propensity score calculation by 

dividing it by peptide length to standardize the pro-

pensity score across the entire peptide.  

 

< 𝛼 >  =
∑ 𝑒𝐻𝑥𝑖𝑖

𝑃𝑒𝑝𝑡𝑖𝑑𝑒 𝐿𝑒𝑛𝑔𝑡ℎ
     (6) 

Hxi represents the ith amino acid’s helix propensity, based 

on the Pace–Schols scale.  

 

CATION-PI INTERACTIONS. Cation-pi interactions 

between the positively charged side chain of an ar-

ginine amino acid residue and the aromatic side 

chain of a tryptophan amino acid residue are of in-

terest because studies suggest that tryptophan con-

taining peptides have higher affinity for deeper in-

sertion into bacterial cell membranes (Hicapie et al, 

2018). Assuming that the peptide adopts an alpha 

helix secondary structure, a cation-pi interaction ex-

ists only for the configuration where an arginine is 

four residues after a tryptophan (Shi et al., 2022). We 

computed a simple count for the number of cation-

pi interactions within a peptide sequence.  

 

𝑇𝑟𝑝 ⟶ 𝐴𝑟𝑔 (𝑖,  𝑖 + 4)    (7) 

 

STATISTICAL ANALYSES AND MACHINE LEARNING TECH-

NIQUES 

SCATTERPLOT MATRIX. Using the psych R pack-

age, a scatterplot matrix was generated to visualize 

the pairwise relationships between the selected bi-

ophysical features. A scatterplot matrix helps to 

graphically and quantitatively identify if multicollin-

earity exists among the various molecular proper-

ties of interest.  

PRINCIPAL COMPONENT ANALYSIS (PCA). As a di-

mensionality reduction technique, PCA is appropri-

ate for our purposes because the DBAASP subset is 

a large data set in which each observation contains 

a high number of features (each peptide is associ-

ated with seven molecular properties). PCA was 

conducted using the factoextra R package and is ac-

complished by linearly transforming the data onto a 

new coordinate system. Information is projected 

onto a two-dimensional space, increasing interpret-

ability while preserving the maximum amount of in-

formation (Jolliffe & Cadima, 2016).  

DECISION TREES. Classification and Regression 

Tree (CART) is a predictive modeling approach. It il-

lustrates how peptide effectiveness (response varia-

ble) can be predicted based on the seven biophysi-

cal properties (explanatory variables). While classifi-

cation trees are used for predicting a qualitative out-

come, regression trees are used for predicting a 

quantitative outcome (Krzywinski & Altman, 2017). 

Using the rpart and rpart.plot R packages, we gen-

erated CART for the DBAASP subset. For the classi-

fication tree, the peptides were separated into two 

classes based on a MIC threshold that we specified: 

noneffective (MIC ≥ 4 µM) and effective (MIC < 4 

µM). For the regression tree, the peptides were par-

titioned based on their biophysical properties, and 

an average log2(MIC) was reported for each sub-

group. log2(MIC) was used in place of the MIC to re-

duce the skewness of the MIC data. log2(MIC) was 

the most sensible transformation to apply because 

MIC assays are based on two-fold serial dilutions.  

RANDOM FORESTS. The random forest algo-

rithm is an ensemble learning method that aggre-

gates multiple decision trees to create more accu-

rate predictions (Brieman, 2001). Like decision 

trees, random forests can handle both regression 

and classification tasks. However, it is a more robust 

approach than CART because it controls overfitting. 

Functions from the randomForest R package were 

implemented to construct a random forest 



ARESTY RUTGERS UNDERGRADUATE RESEARCH JOURNAL, VOLUME I, ISSUE V 

 

 

regression model. The model included all seven bi-

ophysical properties as the explanatory variables 

with log2(MIC) as the response variable. Based on 

the method used to generate the prediction algo-

rithm by Thomas et al. (2010), 70% of the 4,218 pep-

tides in the DBAASP data set were used as the train-

ing set and 30% were used as the testing set. The 

training set was used to build a random forest re-

gression model that would predict the log2(MIC) of 

peptide sequences.  

 

DESIGNING AND TESTING PEPTIDES 

DESIGNING PEPTIDES. Based on the features 

learned from the bioinformatic analyses of the 

DBAASP, we wanted to design new, synthetic pep-

tide sequences. We generated a set of 50,000 ran-

dom peptide sequences that had the same length 

and amino acid distributions as those in our 

DBAASP subset (FIGURE 1). To control the amino acid 

distribution, we assigned each amino acid with a 

probability of occurrence based on the frequency 

that each appeared in the DBAASP subset. We then 

computed the seven biophysical properties of inter-

est for all 50,000 peptide sequences. From there, 

we were able to use the random forest regression 

model to generate predictions for their log2(MIC) 

values. 

 

SELECTING PEPTIDES FOR EXPERIMENTAL TESTING. 

Our goal was to determine if any of the designed 

peptides were effective in vitro against P. aeru-

ginosa infections. Assuming that the peptides with 

the lowest predicted MICs have the greatest likeli-

hoods for being effective, we decided to examine 

20 generated peptide sequences with the lowest 

predicted log2(MIC) values more closely. For the 20 

peptide sequences, we used the Hemolytic Activity 

Prediction for Peptides Employing Neural Networks 

(HAPPENN) tool to calculate the predicted hemo-

lytic activity of the peptides (Timmons & Hewage, 

2020). The HAPPENN tool provides a PROB score 

between 0 and 1. A peptide with a score of 0 is pre-

dicted to be most likely non-hemolytic, while a pep-

tide with a score of 1 is predicted to be most likely 

hemolytic. Because the HAPPENN tool can only 

make predictions on sequences with lengths be-

tween 7 and 35 amino acids long, we were only able 

to deduce the potential hemolytic activity for 11 out 

of the 20 peptides. Of those 11 peptides, 8 peptides 

were predicted to be most likely non-hemolytic. 

Therefore, we decided to select those 8 peptides for 

experimental testing. 

 

 

 
FIGURE 1: (a) Distribution of lengths for n = 4,218 peptides. (b) Distribution of amino acid residues that make up the n = 

4,218 peptides 

  



ARESTY RUTGERS UNDERGRADUATE RESEARCH JOURNAL, VOLUME I, ISSUE V 

 

 

 

Sequence Predicted 𝒍𝒐𝒈𝟐(MIC) 

AKRTQRFPRWKKCLRLGFVGCKGNILKAA -0.175 

RVRKGAGTSRKLKIVKNLGRHIVWFKGIP  0.105 

IRKYRPGLFAKFHKLLKNRKIGGKNL 0.141 

MKSKPTVIMRYRFWRVGKL 0.307 

RQACGTKAKARLRKPRCLTIGRRVVRKFSKWR 0.329 

KMVAKKVKIKRCKRKVHKLPGFGSISIL 0.335 

LIAKYGKHAKFKAKKKPQGISGVPKRFYKALWIG 0.432 

PRRIKTGAAKRKPKLSKKWNQKLLKKLTPFGW  0.464 

 

TABLE 1: List of 8 peptide sequences selected from 50,000 computer generated sequences. These peptides are predicted to 

be both antimicrobial and nonhemolytic. 

 

4 RESULTS AND DISCUSSION  
RELATIONSHIP BETWEEN MOLECULAR PROPERTIES  

The scatterplot matrix illustrates the rela-

tionship between each pair of molecular properties 

(FIGURE 2). Since none of the correlation coefficients 

exceeds 0.70, it is generally agreed upon that none 

of the molecular properties is strongly linearly cor-

related with each other (Ratner, 2009). This demon-

strates that the properties are non-redundant, and 

they may all influence a peptide’s effectiveness. 

However, it is important to note that there exists a 

moderate, positive correlation between hydropho-

bic proportion and maximum average hydrophobic 

residue. This was expected because both hydro-

phobicity-based calculations were dependent on 

Fauchere and Pliska’s scale.  

 

FIGURE 2: Scatterplot Matrix. The diagonal panels provide 

histograms for each molecular property. The panels below 

the diagonal depict bivariate scatter plots, which provide 

a visual representation of the relationship between two 

variables, allowing for pattern identification, outlier detec-

tion, and linear correlation assessment. The panels above 

the diagonal report the Pearson correlation coefficients as-

sociated with the regression line in each scatter plot—a 

value close to +1 or -1 signifies strong correlation while a 

value close to 0 signifies weak correlation. 

 

 

 

 

 
 

 



ARESTY RUTGERS UNDERGRADUATE RESEARCH JOURNAL, VOLUME I, ISSUE V 

 

 

 

The PCA projects information from the 

seven-dimensional space onto orthogonal axes that 

are linear combinations and capture the variation of 

the original variables. The first two dimensions, also 

called principal components, account for 30.1% and 

19.3% (FIGURE 3), respectively, of the total variation in 

the DBAASP subset. This means that the two-dimen-

sional scatter plot, often used in PCA due to its facil-

ity for visual interpretation, is not a faithful represen-

tation of the original data in the seven-dimensional 

space, since it only includes 50% of the variation in 

the DBAASP data set. To account for greater varia-

tion, higher dimensions are needed. Nevertheless, 

the PCA provides insight into how variation is dis-

tributed among the molecular properties. The PCA 

shows that the greatest contributors to the variation 

are length, hydrophobic proportion, and maximum 

average hydrophobic residue.  

 

 
FIGURE 3: Principal Component Analysis. (a) Plot of molec-

ular properties and their respective contributions. (b) Bip-

lot of individual peptide data points and their molecular 

properties. 

 

 
FIGURE 4: CART. For both classification and regression 

trees, each decision node is associated with a condition. If 

the condition is true (or yes) for a peptide, the peptide will 

be sorted down the left branch and either reach a terminal 

node or proceed to the next condition. (a) Classification 

tree built for an n = 4,218 sample, where partitions are 

based on MIC = 4 µM threshold. Each box displays the 

overall classification at that node, and the left value repre-

sents the number of peptides that are effective while the 

right value represents the number of peptides that are not 

effective. (b) Regression tree built for an n = 4,218 sample, 

where the response variable is log2(MIC). Each node box 

displays the predicted log2(MIC) value and the n number 

of peptides that contributed to that predicted value. 

 

RELATIONSHIP BETWEEN MOLECULAR PROPERTIES AND 

PEPTIDE EFFECTIVENESS  

The conditions used in the classification tree 

to split the DBAASP subset indicate that all the mo-

lecular properties, except for cation-pi interactions, 

are important in partitioning the data. For classifica-

tion purposes, we defined a peptide as effective if 



ARESTY RUTGERS UNDERGRADUATE RESEARCH JOURNAL, VOLUME I, ISSUE V 

 

 

its MIC value for P. aeruginosa is less than 4 µM. 

Based on the classification tree that we generated, 

it appears that the lowest ratio of noneffective to ef-

fective peptides (69 to 139) is associated with a 

group of peptides that satisfies the following condi-

tions: net charge ≥ 4.5, maximum average hydro-

phobic moment ≥ 0.6981, and length ≥ 29.5 amino 

acids (FIGURE 4A). Regarding the regression tree that 

we generated, the subgroup of peptides with the 

lowest average log2(MIC) is composed of 208 pep-

tides and satisfies the following conditions: net 

charge ≥ 4.5, maximum average hydrophobic mo-

ment ≥ 0.6977, and length ≥ 29.5 amino acids (FIG-

URE 4B). These conditions are similar to those de-

scribed for the classification tree, suggesting that 

the information provided by the two trees agree. Us-

ing these conditions, we can make predictions for 

new peptide sequences that were not included in 

the DBAASP subset.  

While the decision trees provide simple vis-

ualizations and predictive methods, the predictions 

are not always accurate. By using multiple trees,  

random forests produce more accurate predictions 

and provide more reliable information of variable 

importance. The feature importance attribute of the 

random forest regression model indicates that net 

charge and the maximum average hydrophobic 

moment are the most distinctive biophysical prop-

erties influencing peptide effectiveness (Figure 5). 

The feature importance attribute also confirms that 

the presence of cation-pi interactions contributes 

marginally to the fit of the regression model. The 

other four properties were of similar importance. 

Overall, the random forest regression model had a 

mean of squared residuals of 4.409, which suggests 

that the model is incorrect by 2.1 log2(MIC) units on 

average, and the % variance explained by the model 

is 37.83%. While the fit of the model is lacking, the 

regression model is still worth exploring, as our goal 

is not to be able to predict exact MICs, but to de-

velop a selection method for identifying effective 

peptides based on their biophysical properties.  

 

 

 
FIGURE 5: The importance of a biophysical property is determined by measuring the increase in mean square error and the 

increase in node purity when that property is randomly permuted. Important properties result in substantial changes when 

randomly permuted, leading to greater increases in mean square error and node purity. 

 



ARESTY RUTGERS UNDERGRADUATE RESEARCH JOURNAL, VOLUME I, ISSUE V 

 

 

5 CONCLUSION 
AMPs serve as a rich source of alternatives 

to conventional antibiotics and provide a potential 

solution for the antibiotic resistance crisis. Existing 

AMP databases are continuously being updated to 

account for the expanding knowledge known about 

them. However, a comprehensive understanding of 

the relationship between molecular properties of 

AMPs and their effectiveness against bacterial infec-

tions, especially from a clinical perspective, remains 

elusive. To develop this comprehension and to ac-

celerate the process for screening peptides, this 

study aimed to create a biophysical understanding 

and predictive model for classifying peptides as ef-

fective or noneffective. To do so, we used a relevant 

subset of 4,218 AMPs from the DBAASP.  

To ensure that the selected molecular prop-

erties are not correlated with each other and thus all 

contribute separately to peptide effectiveness, we 

constructed a scatterplot matrix. It is apparent that 

none of the properties are strongly linearly corre-

lated with each other. This was unexpected, consid-

ering that three of the seven selected properties 

were related to the hydrophobicity of the peptide. 

At the same time, this was promising because it in-

dicates that multicollinearity was not present in our 

model, and thus each property individually contrib-

utes to peptide effectiveness. To further examine 

the contributions of each molecular property to the 

total variation within the DBAASP, we performed a 

PCA. Length, hydrophobic proportion, and maxi-

mum average hydrophobic residue had the great-

est amount of information captured in the first two 

dimensions.  

We used CART to gain insight into which bi-

ophysical properties were most important for decid-

ing peptide effectiveness. Both the classification 

tree and regression tree reached a consensus that 

peptides with the highest potential to become effi-

cacious drugs have a maximum average hydropho-

bic moment ≥ 0.75 and have a net charge ≥ 4. The 

finding about net charge is not surprising and is 

consistent with available literature (Jiang et al., 

2008).  

A predictive regression model using the 

random forest algorithm was generated to assess 

the potential antimicrobial activity of a given pep-

tide sequence. This predictive tool was used to eval-

uate a set of 50,000 computer generated peptide 

sequences. Based on the predicted antimicrobial 

and hemolytic activities of those peptide se-

quences, we identified a set of eight peptides that 

may have activity against P. aeruginosa lung infec-

tions. In the future, we wish to assess the accuracy of 

the random forest predictive tool, and we hope to 

confirm the activity of those 8 peptides experimen-

tally by conducting MIC assays. Furthermore, we 

wish to expand upon our understanding of AMPs by 

exploring other methods, including non-linear di-

mensionality reduction, PCA regression, and imple-

menting a more formal search for designing pep-

tides through following a genetic algorithm frame-

work∎  

 

6 ACKNOWLEDGEMENTS 
This work is supported by grant 

#R21AI151746 from the National Institutes of 

Health. 

 

7 REFERENCES 
[1] Ageitos, J. M., Sánchez-Pérez, A., Calo-Mata, P., & 

Villa, T. G. (2017). Antimicrobial peptides (AMPs): An-

cient compounds that represent novel weapons in 

the fight against bacteria. Biochemical pharmacol-

ogy, 133, 117–138.  

HTTPS://DOI.ORG/10.1016/J.BCP.2016.09.018  

[2] Allen, P., Borick, J., & Borick, J. (2020). Acute and 

Chronic Infection Management in CF. Cystic Fibrosis 

in Primary Care, 69–87.  

HTTPS://DOI.ORG/10.1007/978-3-030-25909-9_8 

[3] Breiman, L. (2001). Random Forests. Machine Learn-

ing 45, 5–32. 

HTTPS://DOI.ORG/10.1023/A:1010933404324 

[4] Centers for Disease Control and Prevention. (2019). 

Antibiotic resistant threats in the United States. U.S. 

Department of Health and Human Services. 

HTTPS://WWW.CDC.GOV/DRUGRESISTANCE/PDF/THREATS-RE-

PORT/2019-AR-THREATS-REPORT-508.PDF 

[5] Chen, C.H.; Lu, T.K. (2020). Development and Chal-

lenges of Antimicrobial Peptides for Therapeutic Ap-

plications. Antibiotics, 9, 24.  

[6] Edwards, I. A., Elliott, A. G., Kavanagh, A. M., Zuegg, 

J., Blaskovich, M. A., & Cooper, M. A. (2016). Contri-

bution of Amphipathicity and Hydrophobicity to the 

https://doi.org/10.1016/j.bcp.2016.09.018
https://doi.org/10.1007/978-3-030-25909-9_8
https://doi.org/10.1023/A:1010933404324
https://www.cdc.gov/drugresistance/pdf/threats-report/2019-ar-threats-report-508.pdf
https://www.cdc.gov/drugresistance/pdf/threats-report/2019-ar-threats-report-508.pdf


ARESTY RUTGERS UNDERGRADUATE RESEARCH JOURNAL, VOLUME I, ISSUE V 

 

 

Antimicrobial Activity and Cytotoxicity of β- Hairpin 

Peptides. ACS infectious diseases, 2(6), 442–450. 

HTTPS://DOI.ORG/10.1021/ACSINFECDIS.6B00045 

[7] Eisenberg, D., Weiss, R.M., Terwilliger, T.C., & Wilcox, 

W. (1982). Hydrophobic moments and protein struc-

ture. Faraday Symp. Chem. Soc., 17, 109-120.  

[8] Eisenberg, D., Weiss, R. M., & Terwilliger, T. C. (1984). 

The hydrophobic moment detects periodicity in pro-

tein hydrophobicity. Proceedings of the National 

Academy of Sciences of the United States of America, 

81(1), 140–144.  

HTTPS://DOI.ORG/10.1073/PNAS.81.1.140   

[9] Fauchere, J., Pliska, V (1983). Hydrophobic parame-

ters p of amino-acid side chains from the partitioning 

of N-acetyl-amino-acid amides. Eur. J. Med. Chem, 8, 

369–375.  

[10] Hincapié, O., Giraldo, P., & Orduz, S. (2018). In silico 

design of polycationic antimicrobial peptides active 

against Pseudomonas aeruginosa and Staphylococ-

cus aureus. Antonie van Leeuwenhoek, 111(10), 

1871–1882.  

HTTPS://DOI.ORG/10.1007/S10482-018-1080-2   

[11] Jiang, Z., Vasil, A. I., Hale, J. D., Hancock, R. E., Vasil, 

M. L., & Hodges, R. S. (2008). Effects of net charge and 

the number of positively charged residues on the bi-

ological activity of amphipathic alpha-helical cationic 

antimicrobial peptides. Biopolymers, 90(3), 369–383. 

HTTPS://DOI.ORG/10.1002/BIP.20911  

[12] Jolliffe I.T. and Cadima J. (2016). Principal compo-

nent analysis: a review and recent developments. 

Phil. Trans. R. Soc. A.3742015020220150202. 

HTTPS://DOI.ORG/10.1098/RSTA.2015.0202   

[13] Krzywinski, M., Altman, N. (2017). Classification and 

regression trees. Nat Methods 14, 757–758. 

HTTPS://DOI.ORG/10.1038/NMETH.4370   

[14] Lei, J., Sun, L., Huang, S., Zhu, C., Li, P., He, J., Mackey, 

V., Coy, D. H., & He, Q. (2019). The antimicrobial pep-

tides and their potential clinical applications. Ameri-

can journal of translational research, 11(7), 3919–

3931.  

[15] Li, S., Wang, Y., Xue, Z., Jia, Y., Li, R., He, C., & Chen, 

H. (2021). The structure-mechanism relationship and 

mode of actions of antimicrobial peptides: A review. 

Trends in Food Science & Technology, 109, 103-115. 

DOI:10.1016/J.TIFS.2021.01.005    

[16] Moreau-Marquis, S., Stanton, B. A., & O’Toole, G. A. 

(2008). Pseudomonas aeruginosa biofilm formation 

in the cystic fibrosis airway. Pulmonary Pharmacology 

& Therapeutics, 21(4), 595-599.  

HTTPS://DOI.ORG/10.1016/J.PUPT.2007.12.001     

[17] Pace, C.N., Scholtz, J.M. (1998). A helix propensity 

scale based on experimental studies of peptides and 

proteins. Biophysical Journal 75, 422-427.  

HTTPS://DOI.ORG/10.1016/S0006-3495(98)77529-0   

[18] Pirtskhalava, M., Amstrong, A. A., Grigolava, M., 

Chubinidze, M., Alimbarashvili, E., Vishnepolsky, B., 

Gabrielian, A., Rosenthal, A., Hurt, D. E., & Tartakov-

sky, M. (2021). DBAASP v3: database of antimicro-

bial/cytotoxic activity and structure of peptides as a 

resource for development of new therapeutics. Nu-

cleic acids research, 49(D1), D288–D297. 

HTTPS://DOI.ORG/10.1093/NAR/GKAA991  

[19] Porto, W.F., Irazazabal, L., Alves, E.S.F. et al. (2018). In 

silico optimization of a guava antimicrobial peptide 

enables combinatorial exploration for peptide de-

sign. Nat Commun 9, 1490. doi:  

HTTPS://DOI.ORG/10.1038/S41467-018-03746-3    

[20] Ramazi, S., Mohammadi, N., Allahverdi, A., Khalili, E., 

& Abdolmaleki, P. (2022). A review on antimicrobial 

peptides databases and the computational tools. Da-

tabase: the journal of biological databases and cura-

tion, baac011.  

HTTPS://DOI.ORG/10.1093/DATABASE/BAAC011   

[21] Ratner, B. (2009). The correlation coefficient: Its val-

ues range between +1/−1, or do they?. J Target Meas 

Anal Mark 17, 139–142.  

HTTPS://DOI.ORG/10.1057/JT.2009.5 

[22] Shi, Z., Olson, C. A., & Kallenbach, N. R. (2002). Cat-

ion-pi interaction in model alpha-helical peptides. 

Journal of the American Chemical Society, 124(13), 

3284–3291.  

HTTPS://DOI.ORG/10.1021/JA0174938   

[23] Thomas, S., Karnik, S., Barai, R. S., Jayaraman, V. K., & 

Idicula-Thomas, S. (2010). CAMP: a useful resource 

for research on antimicrobial peptides. Nucleic acids 

research, 38(Database issue), D774–D780.  

HTTPS://DOI.ORG/10.1093/NAR/GKP1021  

[24] Timmons, P. B., & Hewage, C. M. (2020). HAPPENN is 

a novel tool for hemolytic activity prediction for ther-

apeutic peptides which employs neural networks. 

Scientific reports, 10(1), 10869.  

HTTPS://DOI.ORG/10.1038/S41598-020-67701-3  

[25] Ye, Z., & Aparicio, C. (2019). Modulation of supramo-

lecular self-assembly of an antimicrobial designer 

peptide by single amino acid substitution: Implica-

tions on peptide activity. Nanoscale advances, 1(12), 

4679–4682.  

HTTPS://DOI.ORG/10.1039/C9NA00498J   

  

https://doi.org/10.1021/acsinfecdis.6b00045
https://doi.org/10.1073/pnas.81.1.140
https://doi.org/10.1007/s10482-018-1080-2
https://doi.org/10.1002/bip.20911
https://doi.org/10.1098/rsta.2015.0202
https://doi.org/10.1038/nmeth.4370
https://doi.org/10.1016/j.pupt.2007.12.001
https://doi.org/10.1016/S0006-3495(98)77529-0
https://doi.org/10.1093/nar/gkaa991
https://doi.org/10.1038/s41467-018-03746-3
https://doi.org/10.1093/database/baac011
https://doi.org/10.1057/jt.2009.5
https://doi.org/10.1021/ja0174938
https://doi.org/10.1093/nar/gkp1021
https://doi.org/10.1038/s41598-020-67701-3
https://doi.org/10.1039/c9na00498j


ARESTY RUTGERS UNDERGRADUATE RESEARCH JOURNAL, VOLUME I, ISSUE V 

 

 

Alisha Zhu is a student in the Honors College at Rutgers University – New Bruns-

wick. She is majoring in Biomedical Engineering and minoring in Statistics. 

Over the past three years, she has been working as an undergraduate re-

searcher in Dr. Charles Roth’s Lab, which she first joined through the Aresty 

Research Assistant Program. Her research focuses on using computational sta-

tistics and machine learning techniques to study the relationship between the 

biophysical properties of antimicrobial peptides and their drug efficacies. Out-

side of research, she is the President of the Society of Women Engineers (SWE) 

and is the External Vice President of the Chinese Student Organization (CSO).  

 

Contact: ALISHA.ZHU@RUTGERS.EDU 


