From owner-chemistry@ccl.net Wed Oct 27 08:44:00 2010 From: "Andras Borosy andras.borosy::givaudan.com" To: CCL Subject: CCL: Descriptors Message-Id: <-43012-101027060202-7547-8MtSGoS+eigd+zJ5jcvGYQ^^^server.ccl.net> X-Original-From: Andras Borosy Content-Type: multipart/alternative; boundary="=_alternative 003716ADC12577C9_=" Date: Wed, 27 Oct 2010 12:01:50 +0200 MIME-Version: 1.0 Sent to CCL by: Andras Borosy [andras.borosy:+:givaudan.com] This is a multipart message in MIME format. --=_alternative 003716ADC12577C9_= Content-Type: text/plain; charset="ISO-8859-1" Content-Transfer-Encoding: quoted-printable Dear George, If I am not mistaken, you have MOE. Well, then use the ga.svl! It is very=20 effective genetic algorithm and selects several good equations. Generally=20 the first one (with least LOF) is the best one. On the other hand you must = know the experimental error of your dependent variable (Y values,=20 properties, activities), before you would start any QSPR building!=20 Best wishes, Dr. Andr=E1s P=E9ter Borosy Scientific Modelling Expert Fragrance Research Givaudan Schweiz AG - Ueberlandstrasse 138 - CH-8600 - D=FCbendorf -= =20 Switzerland T:+41-44-824 2164 - F:+41-44-8242926 - http://www.givaudan.com "Erik-Jan Ras Erik-Jan.Ras..avantium.com" =20 Sent by: owner-chemistry+andras.borosy=3D=3Dgivaudan.com[-]ccl.net 26.10.2010 20:23 Please respond to "CCL Subscribers" To "Borosy, Andras " cc Subject CCL: Descriptors Sent to CCL by: Erik-Jan Ras [Erik-Jan.Ras~~avantium.com] Dear George, As already indicated by others, there is no uniform selection method for=20 choosing which descriptors to use. Some guidelines, depending on the=20 modeling method you use may still be helpfull. If you're using PLS models, a good starting point is the variable=20 importance (VIP) for each of the variables in your model. A variable with=20 a high VIP will have a high impact on your model performance. Typically=20 you start your modeling exercise with all available variables. After that, = in small iterative steps, you reduce your model. At each stage you have to = carefully evaluate predictive power of your model. Ideally you would use a = substantially large external validation set to assess predictive power. Also keep in mind the fact that per response (Y) in theory only one latent = variable should be required in your model. If (many) more latent variables = are required you're dealing with variations in your descriptor space (X)=20 that are orthogonal (uncorrelated) to your response (Y). In this case you=20 may want to consider using OPLS in stead of PLS. Generally speaking, these methods are implemented in commercial packages=20 like Simca-P and work quite well (also pretty well documented and=20 referenced). With a bit more effort in environments like Matlab, Scilab or = R many open source libraries are available as well. Regards, Erik-Jan =5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F= =5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F > From: owner-chemistry+erikjan.ras=3D=3Davantium.com=5F-=5Fccl.net=20 [owner-chemistry+erikjan.ras=3D=3Davantium.com=5F-=5Fccl.net] On Behalf Of = George=20 Lawrence geoe2##hotmail.com [owner-chemistry=5F-=5Fccl.net] Sent: Tuesday, October 26, 2010 12:51 PM To: Erik-Jan Ras Subject: CCL: Descriptors Sent to CCL by: "George Lawrence" [geoe2%hotmail.com] While building a model for a set of compounds, how does one make the=20 choice of molecular descriptors, I am using MOE which has about 333=20 different descriptors. I noticed that some have the same suffix or prefix. For example: GCUT (could be SlogP, SMR or PEOE) and then there is SlogP=5F = vsa, SMR=5F vsa, PEOE=5Fvsa which have different numbers attach to them. Wh= at=20 does this mean? Do they describe the same thing? How does the numbers relate to each=20 descriptor? What are the best methods to use to decide the right choice of=20 descriptors? George Lawrence Geoe2[a]hotmail.com Kent=20 U.K.http://www.ccl.net/cgi-bin/ccl/send=5Fccl=5Fmessagehttp://www.ccl.net/c= hemistry/sub=5Funsub.shtmlhttp://www.ccl.net/spammers.txtThis=20 email (including its attached files and other content) is confidential and = intended only for the use by named addressee. Unauthorized use,=20 dissemination, disclosure and/or copying are prohibited. This email,=20 attachments and (any part of) its content are (1) intended for the named=20 addressee(s) only, and (2) strictly confidential and proprietary. All=20 rights are reserved by Avantium Holding B.V. and its subsidiaries=20 ('Avantium'). Any unauthorized use, dissemination, disclosure and/or=20 copying is strictly prohibited, except after prior and express written=20 permission by Avantium. Avantium is not responsible for the correct=20 transmission and timely receipt of this email and its content. Should you=20 have received this email, attachments and its content by mistake, please=20 bring this to our attention and destroy this email in full. Thank you.=20 http://www! .avantium.com/about/legal-disclaimer/ -=3D This is automatically added to each message by the mailing script =3D-http://www.ccl.net/cgi-bin/ccl/send=5Fccl=5Fmessagehttp://www.ccl.net/cgi-bin/ccl/send=5Fccl=5Fmessage Subscribe/Unsubscribe:=20 http://www.ccl.net/chemistry/sub=5Funsub.shtmlJob: http://www.ccl.net/jobs=20http://www.ccl.net/spammers.txt--=_alternative 003716ADC12577C9_= Content-Type: text/html; charset="ISO-8859-1" Content-Transfer-Encoding: quoted-printable
Dear George,

If I am not mistaken, you have MOE. Well, then use the ga.svl! It is very effective genetic algorithm and selec= ts several good equations. Generally the first one (with least LOF) is the best one. On the other hand you must know the experimental error of your dependent variable (Y values, properties, activities), before you would start any QSPR building!

Best wishes,

Dr. Andr=E1s P=E9ter Borosy
Scientific Modelling Expert

Fragrance Research
Givaudan Schweiz AG  -  Ueberlandstrasse 138  -  CH-8600  -  D=FCbendorf  -  Switzerland
T:+41-44-824 2164  -  F:+41-44-8242926    -  http:= //www.givaudan.com




"Erik-Jan Ras Er= ik-Jan.Ras..avantium.com" <owner-chemistry[-]ccl.net>
Sent by: owner-chemistry+andras.boro= sy=3D=3Dgivaudan.com[-]ccl.net

26.10.2010 20:23
Please respond to
"CCL Subscribers" <chemistry[-]ccl.net>

To
"Borosy, Andras " <andras.borosy[-]givaudan.com>
cc
Subject
CCL: Descriptors






Sent to CCL by: Erik-Jan Ras [Erik-Jan.Ras~~avantium.com]
Dear George,

As already indicated by others, there is no uniform selection method for choosing which descriptors to use. Some guidelines, depending on the modeli= ng method you use may still be helpfull.

If you're using PLS models, a good starting point is the variable importance (VIP) for each of the variables in your model. A variable with a high VIP will have a high impact on your model performance. Typically you start your modeling exercise with all available variables. After that, in small iterative steps, you reduce your model. At each stage you have to carefully evaluate predictive power of your model. Ideally you would use a substantia= lly large external validation set to assess predictive power.

Also keep in mind the fact that per response (Y) in theory only one latent variable should be required in your model. If (many) more latent variables are required you're dealing with variations in your descriptor space (X) that are orthogonal (uncorrelated) to your response (Y). In this case you may want to consider using OPLS in stead of PLS.

Generally speaking, these methods are implemented in commercial packages like Simca-P and work quite well (also pretty well documented and reference= d). With a bit more effort in environments like Matlab, Scilab or R many open source libraries are available as well.

Regards,
Erik-Jan


=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F= =5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F=5F
> From: owner-chemistry+erikjan.ras=3D=3Davantium.com=5F-=5Fccl.net [own= er-chemistry+erikjan.ras=3D=3Davantium.com=5F-=5Fccl.net] On Behalf Of George Lawrence geoe2##hotmail.com [owner-chemistry=5F-=5Fccl.= net]
Sent: Tuesday, October 26, 2010 12:51 PM
To: Erik-Jan Ras
Subject: CCL: Descriptors

Sent to CCL by: "George  Lawrence" [geoe2%hotmail.com]
While building a model for a set of compounds, how does one make the choice of molecular descriptors, I am using MOE which has about 333 different descriptors. I noticed that some have the same suffix or prefix.
For example: GCUT (could be SlogP, SMR or PEOE) and then there is SlogP=5F vsa, SMR=5F vsa, PEOE=5Fvsa which have different numbers attach to them. Wh= at does this mean?

Do they describe the same thing? How does the numbers relate to each descr= iptor?
What are the best methods to use to decide the right choice of descriptors?=

George Lawrence
Geoe2[a]hotmail.com
Kent U.K.http://www.ccl.net/cgi-bin/ccl/send=5Fccl=5Fmessagehttp://www.ccl.= net/chemistry/sub=5Funsub.shtmlhttp://www.ccl.net/spammers.txtThis email (including its attached files and other content) is confidential and intended only for the use by named addressee. Unauthorized use, dissemi= nation, disclosure and/or copying are prohibited. This email, attachments and (any part of) its content are (1) intended for the named addressee(s) only, and (2) strictly confidential and proprietary. All rights are reserved by Avantium Holding B.V. and its subsidiaries ('Avantium'). Any unauthorized use, dissemination, disclosure and/or copying is strictly prohibited, except after prior and express written permission by Avantium. Avantium is not responsible for the correct transmission and timely receipt of this email and its content. Should you have received this email, attachments and its content by mistake, please bring this to our attention and destroy this email in full. Thank you. http://www!
.avantium.com/about/legal-disclaimer/



-=3D This is automatically added to each message by the mailing script =3D-=      http://www.ccl.net/cgi-bin/ccl/send=5Fccl=5Fmessage
     http://www.ccl.net/cgi-bin/ccl/send=5Fccl=5Fmessage
     http://www.ccl.net/chemistry/sub=5Funsub.shtml
     http://www.ccl.net/spammers.txt



--=_alternative 003716ADC12577C9_=--