File(s) under permanent embargo
Performance comparison of partial least squares-related variable selection methods for quantitative structure retention relationships modelling of retention times in reversed-phase liquid chromatography
journal contribution
posted on 2015-12-11, 00:00 authored by Mohammad Talebi, Georg Schuster, Robert ShellieRobert Shellie, Roman Szucs, Paul R HaddadThe relative performance of six multivariate data analysis methods derived from or combined with partial least squares (PLS) has been compared in the context of quantitative structure-retention relationships (QSRR). These methods include, GA (genetic algorithm)-PLS, Monte Carlo uninformative variable elimination (MC-UVE), competitive adaptive reweighted sampling (CARS), iteratively retaining informative variables (IRIV), variable iterative space shrinkage approach (VISSA) and PLS with automated backward selection of predictors (autoPLS). A set of 825 molecular descriptors was computed for 86 suspected sports doping compounds and used for predicting their gradient retention times in reversed-phase liquid chromatography (RPLC). The correlation between molecular descriptors selected by each technique and the retention time was established using the PLS method. All models derived from a selected subset of descriptors outperformed the reference PLS model derived from all descriptors, with very small demands of computational time and effort. A performance comparison indicated great diversity of these methods in selecting the most relevant molecular descriptors, ranging from 28 for CARS to 263 for MC-UVE. While VISSA provided the lowest degree of over-fitting for the training set, CARS demonstrated the best compromise between the prediction accuracy and the number of selected descriptors, with the prediction error of as low as 46s for the external test set. Only ten descriptors were found to be common for all models, with the characteristics of these descriptors being representative of the retention mechanism in RPLC.
History
Journal
Journal of Chromatography AVolume
1424Pagination
69 - 76Publisher
ElsevierLocation
Amsterdam, The NetherlandsPublisher DOI
ISSN
0021-9673eISSN
1873-3778Language
engPublication classification
C1 Refereed article in a scholarly journalCopyright notice
2015, ElsevierUsage metrics
Categories
No categories selectedKeywords
Genetic algorithm (GA)Molecular descriptorsPartial least squares (PLS)QSRRRPLCRetention time predictionAlgorithmsChromatography, Reverse-PhaseDoping in SportsLeast-Squares AnalysisModels, TheoreticalMonte Carlo MethodMultivariate AnalysisQuantitative Structure-Activity RelationshipScience & TechnologyLife Sciences & BiomedicinePhysical SciencesBiochemical Research MethodsChemistry, AnalyticalBiochemistry & Molecular BiologyChemistryMEGAVARIATE ANALYSISPLS-REGRESSIONPREDICTIONVALIDATIONSTRATEGYTOOL
Licence
Exports
RefWorks
BibTeX
Ref. manager
Endnote
DataCite
NLM
DC