Journal Article
ARTICLE IN PRESS
0940-2993/$ - se
doi:10.1016/j.et
$ Presented
Inhalation, Tox � Correspond
E-mail addr
Experimental and Toxicologic Pathology 60 (2008) 213–224
www.elsevier.de/etp
Statistical analysis of in vitro data for risk assessment – Exemplified
for a case of Ames test data $
Markus Roller a,�
, Michaela Aufderheide b
a Advisory Office for Risk Assessment, Doldenweg 14, D-44229 Dortmund, Germany b Fraunhofer Institute of Toxicology and Experimental Medicine (ITEM), Nikolai-Fuchs-Str. 1, D-30625 Hannover, Germany
Received 4 December 2007; accepted 30 January 2008
Abstract
Ames test data of experiments with smoke of six cigarette types were used for dose–response analysis and for derivation of a measure of mutagenic potency. Each cigarette type had been tested using a smoking machine and four dilutions of the smoke of each of seven cycles (one to seven cigarettes). Three plates had been exposed per cigarette number/smoke dilution combination and three control plates had been simultaneously exposed to clean air with each set of smoke-exposed plates. It was the aim of the statistical analysis to determine the slopes of dose–response relationships of various cigarette types and to compare them using statistical tests. Basically, the following procedure is recommended: (1) calculate a dose measure on the basis of the number of smoked cigarettes per cycle and dilution air flow. (2) Use the absolute count values of the individual plates as effect variable. (3) Describe the dose–response relations of the individual cigarette types on the basis of all available data with a polynomial model by means of Poisson regression analysis accounting for overdispersion. (4) Identify the linear dose–response region using the likelihood ratio test and restrict the data set to this region. (5) Use the slope of the linear model in the restricted data set as the basis of the mutagenicity measure. (6) Compare the slope for the individual cigarette type with the slope for a reference cigarette by means of multivariate Poisson regression using the likelihood ratio test and accounting for overdispersion. It is finally recommended to express the mutagenic potency as percentages related to the reference cigarette K2R4F. This type of cigarette was set here equal to 100%; the following values are then obtained for some commercially available cigarette types: type A 25%, type B 90%, type C 119%, type D 13%, type E 59%. The differences are statistically significant. r 2008 Elsevier GmbH. All rights reserved.
Keywords: Cigarette smoke; Ames test; Dose–response relationship; Poisson regression; Overdispersion; Mutagenic potency
Introduction
It is one of the tasks of regulatory toxicology to find out whether a chemical has carcinogenic or mutagenic
e front matter r 2008 Elsevier GmbH. All rights reserved.
p.2008.01.008
at the Congress on Alternative Test Methods in
icology, Berlin, Germany, 7–9 May 2007.
ing author. Tel./fax: +49 231 79 79 489.
ess: [email protected] (M. Roller).
potential, this is called hazard analysis. Recent risk analysis processes are more interested in the ‘‘potency’’ of a substance, where the term potency is used for a quantitative measure as opposed to potential or hazard, which just refers to a property or quality of a substance. Analysis of potency is the basis for quantitative risk assessment and analysis of dose–response relationships is the basis for the calculation of potency. Therefore, it is important to use adequate statistical methods for the
ARTICLE IN PRESS
Table 1. Characteristics of test cigarettes
Cigarette
type
mg tar/
cigarette
mg nicotine /
cigarette
Type of filter
A 3 0.3 Charcoal
B 8 0.7 Charcoal
C 10 0.8 Charcoal
D 1 0.1 Cellulose
acetate
E 4 0.4 Cellulose
acetate
M. Roller, M. Aufderheide / Experimental and Toxicologic Pathology 60 (2008) 213–224214
analysis of dose–response relationships both for in vivo and for in vitro data. When dose–response relationships of long-term carcinogenicity studies are analysed, a measure of carcinogenic potency can be obtained as a result; and the potencies of various substances can be compared. The key word for such analyses is ‘‘regression analysis’’. It has become the methodological standard to use the so-called maximum likelihood method for regression analysis of the data from carcinogenicity studies. Modules for such analyses are implemented in many commercial software packages; a user-friendly version, the Benchmark Dose Software (BMDS), is commonly available on the website of the US Environ- mental Protection Agency (US EPA, 2007). The maximum likelihood method is described in the Help- file of the BMDS software as follows: ‘‘Maximum likelihood is the process of estimating the model
parameters such that a likelihood function is maximized according to the data. In other words, parameter values
are ‘chosen’ such that the subject model (y) obtains the best possible fit to the data, given the constraints of the
model’s parameter structure’’. An analysis of this type, using the multistage model, has recently been applied on a large carcinogenicity study with 16 different dusts of respirable granular bio-durable particles without known
significant specific toxicity. The analysis showed that the carcinogenic potency of the dusts is dependent on retained volume and on particle size, e.g. the potency of nanoparticles was five to six times higher than the potency of fine particles with a mean diameter of 1.8–4 mm (Roller and Pott, 2006).
By using analogous methods, potency can also be analysed from in vitro data, similar to the statistical analysis of in vivo data. However, relatively simple methods were often used in the past for the analysis of dose–response relationships of in vitro studies. For example, a comparison of the genotoxic effects of various fibrous dusts in an in vitro model based on anchorage independent growth was carried out using the so-called method of least squares (Riebe-Imre et al., 1994). It is not the subject of this paper to explain the details of various regression techniques like the method of least squares and the maximum likelihood method in general. It should just be mentioned that use of the method of least squares – this is the kind of analysis implemented in general spreadsheet programs and in some pocket calculators – can be criticized, because it may not adequately take into account the statistical distribution of the data and therefore particularly the results of statistical tests may be misleading. This means, a statistical test may indicate ‘‘significance’’, but this evaluation may not be justified because the statistical test used was not adequate.
This paper deals with the application of the maximum likelihood method on regression analysis of Ames test data. It is the aim of the analyses to derive a ‘‘mutagenicity index’’ which allows the differentiation of
various cigarette types according to the mutagenic potency of their smoke in a specially designed Ames test.
Experimental materials and methods
A smoking machine was used for experimental investigations of cigarette smoke with the Ames test allowing exposure of a large number of plates. The test cigarettes were purchased on the open market. They all have King Size format and contain cased flavoured tobacco blends. They are tip ventilated, those with relatively low tar content rather highly. Three test cigarettes have charcoal filters, the remaining two have cellulose actetate filters. The levels of tar and nicotine per cigarette, as printed on the packs, are shown in Table 1. The smoke of the test cigarettes was compared to the standard laboratory cigarette K2R4F. Each cigarette type was examined in four dilutions of the smoke and this was done in each of seven cycles, i.e. with one smoked cigarette and with up to seven smoked cigarettes. Exposures were done with the following four flow rates of synthetic air, which served for the dilution of the smoke: 1.5–2.0–2.5–3.0 L/min. Each dilution air flow was usually applied in altogether three repetitions of the experiment. From the combinations of smoke dilution and number of smoked cigarettes, many different exposure strengths resulted, which could be used in terms of dose–response relations. Three plates were exposed at the same time for each exposure strength; likewise at the same time three plates were exposed to the synthetic air which was used for the production of the smoke dilutions. Details of the methods were described by Aufderheide and Gressmann (2007). For dose–response analyses, firstly the dose and response variables had to be defined.
Statistical analysis
Selection of dose variable
Common dose measures in toxicological investiga- tions with exposure via the air are the concentration of
ARTICLE IN PRESS M. Roller, M. Aufderheide / Experimental and Toxicologic Pathology 60 (2008) 213–224 215
a test substance in air and the cumulative exposure. The cumulative exposure is the mathematical product of (mean) substance concentration and exposure duration. These dose measures could not be used here, because no indicator substance was defined and/or measured. It was exactly the aim of the investigations to characterize the mutagenicity of smoke independently of the definition of an individual indicator or reference substance. A dose measure had therefore to be formed from the number of smoked cigarettes per single experiment and from the air volumes and/or air flows.
The smoke of a so-called puff enters the exposure system from a standardized volume of 35 mL during a period of 7 s. The volume of the exposure system itself is relatively small. During an experiment, synthetic air is continuously pumped through the system with a flow rate of 1.5 or 2.0 or 2.5 or 3.0 L/min. A puff takes place every 60 s. The calculation of the ‘‘dilution’’ and thus the ‘‘dose’’ depends on the assumption of which period is considered as relevant for the actual exposure of the plates. It does not appear plausible that smoke components remain effective in the system over the entire 60 s between the puffs. Therefore, we preferred the calculation of the smoke dilution on the basis of the 7-s puff expulsion duration. The dose measure ‘‘cigarette equivalents’’ (cig.equiv.) may also be termed ‘‘cigarettes/ dilution’’. This quantity is dimensionless and is calcu- lated from the number of the cigarettes smoked per cycle and from the dilution of the smoke due to the puff expulsion rate of 35 mL/7 s and the flow rate of the synthetic air. For example, a dilution of 35 mL/210 mL (175+35 mL) results for one smoked cigarette and for a dilution air flow of 1.5 L/min (175 mL/7 s); in this case, this yields a dilution of 1/6 equal to 0.17 cigarette equivalents. If two cigarettes are smoked per cycle using the same smoke dilution, a dose of 0.33 cigarette equivalents follows. Smoking one cigarette per cycle with a dilution air flow of 2.0 L/min gives a dilution of 1/7.7 and thus corresponds to 0.13 cigarette equivalents; and so on.
Selection of response variable
The primary datum obtained in the Ames test is the number of the revertant colonies per plate. If statistical tests are to be carried out regarding the question of statistical significance of some dose–response parameters, then the type of regression analysis depends on the statistical distribution of the effect variable. Accordingly, some statistical literature dealing with the analysis of Ames test data concentrates on the question of the type of statistical distribution of the variable ‘‘revertant counts per plate’’ and on the adequate regression analysis for this quantity (Margolin et al., 1981; Myers et al., 1981; Bernstein et al., 1982; Kim and Margolin, 1999).
Traditionally, the results of the colony counting of the exposed groups are given as increases relative to the counts in the control groups (relative counts or count ratio). The two approaches – absolute count numbers versus relative increase related to controls – are not equivalent; they can lead to different interpretations of the results. An optimal way to a decision for one of the two approaches could be the exact knowledge of the biological mechanisms or knowing the clear answer to questions like ‘‘will DNA damages by the test sub- stances contribute to background DNA damage in an additive or in a multiplicative manner?’’ or ‘‘what are cause and effect relationships between the sensitivity of a cell, background DNA damage and DNA damage induced by the test substances?’’. The unequivocal answer to these questions is not known to us.
The difference between absolute and relative counts is important especially if there are large differences between the control counts of several experiments. Table 2 shows, however, that the differences between the control groups of the experiments are relatively small and that therefore the impact of the selection of an absolute or relative effect variable will be small. For each cigarette type, about 250 control plates were used. The overall median of all 1500 control plates amounts to 20, the median values of the controls of the individual cigarette types differ only by maximally one count from the overall median, i.e. there are only the three median values 19, 20 and 21. Because of the relatively small variability of the revertant counts of controls, an important influence of methodology ‘‘absolute counts’’ versus ‘‘relative counts’’ on the final result is unlikely. Also, preliminary regression analyses using a polynomial model and the least squares method showed better model fit when using the absolute effect variable ‘‘absolute number of revertant colonies per plate’’ compared to using the relative effect variable ‘‘increase of the ratio of revertant colonies per plate related to controls’’. Altogether, it can be concluded that for the planned regression analyses the use of relative count values does not offer advantages. In accordance with the literature (e.g. Kim and Margolin, 1999), we therefore used the raw data (absolute count number per plate) as effect variable.
Regression analyses under consideration of the
statistical distribution of the effect variable
It is emphasized in the fine review of Kim and Margolin (1999) over statistical methods for the evaluation of dose–response relations of Ames test data that the kind of regression analysis should depend on the kind of the statistical distribution of the effect variable. Due to theoretical considerations, the revertant counts of the Ames test have been regarded as Poisson- distributed. A regression analysis according to the
ARTICLE IN PRESS
Fig. 1. Frequency distribution of all (pooled) control counts
(exposed to clean air only); sample size about 1500.
Comparison of the observed frequencies with expected curves
(GP distribution ¼ generalized Poisson distribution according
to Consul and Jain, 1973).
Table 2. Basic descriptive statistics of the revertant counts of the controls (exposed to synthetic air only)
Cigarette type (to which the
controls were assigned)
No. of control
plates
Revertants per control plate
Median Arithmetic
mean
Variance S.D.
K2R4F 246 21 21.22 36.99 6.08
Type A 252 20 20.48 40.52 6.37
Type B 249 20 20.74 53.92 7.34
Type C 252 19 19.67 40.72 6.38
Type D 252 19 18.97 28.27 5.32
Type E 252 20 20.21 35.39 5.95
Total 1503 20 20.21 39.68 6.30
M. Roller, M. Aufderheide / Experimental and Toxicologic Pathology 60 (2008) 213–224216
maximum likelihood method on the basis of the Poisson distribution is called Poisson regression. The Poisson distribution is characterized completely by one para- meter; this parameter is at the same time mean and variance. Estimated mean and variance of a sample of a Poisson-distributed variable are therefore approxi- mately equal. It has been reported in the literature that the calculated variance of Ames counts is frequently larger than the mean. This is referred to as over- dispersion or extra-Poisson or hyper-Poisson variability.
Therefore, some further distributions have been proposed to account for this hyper-Poisson variability. These distributions may be summarized as generalized Poisson distributions. For example, Bernstein et al. (1982) refer to some earlier work of Consul and Jain (1973) who describe a family of generalized Poisson distributions. According to the considerations of Bern- stein et al. (1982), the following general equation for the probability function P(x) may be formulated, where x is the number of observed counts per plate, m is the mean and k is an additional parameter:
PðxÞ¼ ðm=
ffiffiffi
k p Þððm=
ffiffiffi
k p Þþð1�ð1=
ffiffiffi
k p ÞÞxÞ
x�1 e�ððm=
ffiffi
k p Þþð1�ð1=
ffiffi
k p ÞÞxÞ
x! .
(1)
This function is defined for k greater than 1, and variance is
V ¼ km. (2)
The parameter k may therefore also be referred to as dispersion factor. For k ¼ 1, Eq. (1) reduces to the Poisson distribution.
Statistical distribution of the effect variable
A large data base for the analysis of the distribution of our Ames test data is provided by the data of the control plates. Fig. 1 shows the distribution of all pooled control counts, which results in a sample size of approximately 1500 observations. By using the mean and variance of the pooled data set (see Table 2), the expected distribution curves according to a Poisson and
to a generalized Poisson (GP) distribution were calcu- lated (m ¼ 20.2, V ¼ 39.7, and according to Eq. (2): k ¼ 1.96, respectively). After visual inspection of the picture, both theoretical distributions seem to fit the observations somehow, though the fit is anything else but perfect. According to a Chi
2 -test, the deviation of
the data from the expected distributions is highly significant. In our opinion, this significant deviation should not be overestimated, because empirical data will always follow mathematical distribution formulae only approximately. On the basis of the empirical data, no clear differentiation between the distribution forms can be made. Bernstein et al. (1982) came to a similar conclusion.
Regression models for the complete data sets
Two decisions have to be made for the regression analysis: firstly, the kind of regression procedure, and secondly, the regression function (model). In their review paper, Kim and Margolin (1999) have described
ARTICLE IN PRESS M. Roller, M. Aufderheide / Experimental and Toxicologic Pathology 60 (2008) 213–224 217
several so-called biologically based mechanistic models. All these models go back to similar theoretical considerations on mutagenic and lethal ‘‘hits’’. The paper also refers to a non-commercial computer program, which was named SALM. The model in SALM can be formulated as follows (where b0, b1 and b2 are the parameters to be estimated by regression analysis):
counts ¼ðb0 þ b1 �doseÞ� ð2� e b2�doseÞ. (3)
In the following, we refer to this model as the ‘‘Margolin model’’. This model contains a linear term, which characterizes the lower region of the dose–response relationship, and a further term, which describes a dying of cells due to substance toxicity – and the associated decrease of the revertant frequency with higher dosages. For comparison with a simple polynomial model, we implemented the Margolin model (Eq. (3)) within the module ‘‘Nonlinear Regression’’ of the commercial software STATISTICA, which allows both the regres- sion function and the so-called loss function L to be defined by the user (StatSoft, 2000). To obtain maximum likelihood estimates according to Poisson
Fig. 2. Dose–response relationships for the smoke of six cigarette ty
were used for three types of regression analysis (see text). The indiv
analysis; the data are presented here in terms of means and standar
regression, we defined the following loss function: L ¼ PRED�OBS log(PRED) (where OBS are the va- lues of the observed counts and PRED are predicted values according to the regression results; log is the natural logarithm).
Fig. 2 contains three curves per plot (i.e. for each of the six cigarette types): the ‘‘biologically based’’ Margolin model and two curves of a polynomial model. The ‘‘polynomial’’ is limited thereby to the following simple linear-quadratic form (with parameters b0, b1 and b2 to be estimated):
counts ¼ b0 þ b1 �doseþ b2 �dose 2 . (4)
This model has been fitted to the data by two different techniques: firstly, by the common method of least squares, and secondly, by means of Poisson regression (module ‘‘Generalized Linear Model’’ in STATISTICA), i.e. the model equation (Eq. (4)) is the same for the two curves in each plot, but the procedures to fit the function to the data – and thus the estimated values of the parameters – are different. The curves of all three regression analyses per plot are very similar; they can hardly be distinguished visually.
pes on the basis of the complete Ames test data sets. The data
idual count data per plate (raw data) were used for regression
d deviations, this is for illustrative reasons only.
ARTICLE IN PRESS M. Roller, M. Aufderheide / Experimental and Toxicologic Pathology 60 (2008) 213–224218
Because the curves in all models are so similar and because the polynomial is much easier to handle than the Margolin model, we preferred the polynomial (linear-quadratic) model for the following analyses. This is particularly justified since for the final characteriza- tion of the mutagenic potency per cigarette and for their statistical comparison only the linear region of the dose–response curve will be used (see below). The problem that remains is overdispersion. As can be seen from Fig. 2, it is not likely that the method of regression analysis has great influence on the ‘‘best fit’’ results, but it will certainly influence estimated variances and the results of statistical tests with regard to significance of slope values. However, there is a quite useful method to account for overdispersion in Poisson regression, which is implemented in STATISTICA (StatSoft, 2000). This method is described in the next section. We found that this method produced results similar to a more complicated use of the generalized Poisson distribution according to Eq. (1). This combination of the usual Poisson regression with a dispersion factor is technically more feasible than some other possibly adequate methods and it does not require defining any specific generalized Poisson distribution.
Identification of the linear region
It is our aim to compare the slopes of the dose–response relationships of different aerosols statis- tically. This is problematic if there are two or more parameters in a model which contribute to the curvature. In the linear-quadratic model, for example, some part of the increase of the effect with dose may be contained in the linear parameter and another part in the quadratic term. For one cigarette type, the parameter estimate of the linear term may have a large value and the estimate for the quadratic term may be small; for a second cigarette type, the opposite may be true. It may be very difficult in such a situation to assign the higher mutagenic potency to one or the other of these cigarette types. The same problem is encountered with the Margolin model (Eq. (3)), which also has two parameters (b1, b2). A solution to this problem is to base the statistical comparison of two dose–response rela- tionships on the linear region of the curve. Then just one slope parameter per cigarette type has to be evaluated. This method appears appropriate because it is evident that all dose–response relations examined here do have a relatively large linear region (Fig. 2). We then need a method to define the linear region in a reproducible manner.
This problem is also discussed in the literature. Bernstein et al. (1982) describe an algorithm, which is based on an iterative fitting of a model to the data and successive elimination of the highest dose in each iteration. The algorithm preferred by Bernstein et al. (1982) is quite complex. The following algorithm is
similar to the method of Bernstein et al. (1982) – and has also been taken into account by them – but it appears clearer and easier. Bernstein et al. (1982) assumed that such an algorithm yields results similar to their preferred method:
1.
Fit a linear-quadratic model to the data by means of Poisson regression accounting for overdispersion.
2.
If the estimate for the quadratic term is negative and if it contributes significantly to the regression, then exclude the highest dose from the further computa- tions and repeat step 1.
3.
If the estimate for the quadratic term does not contribute significantly to the regression, then apply Poisson regression with the linear model and use the estimate of the regression coefficient (slope) as a measure of mutagenic potency.
Two things have to be explained in greater detail here: ‘‘significant contribution to the regression’’ and ‘‘ac- counting for overdispersion’’. Significant contribution of a term to the regression may be tested by means of the so-called likelihood ratio test. In our case, we may calculate the difference of the log-likelihood from the regression with and without the quadratic term; the ‘‘significance’’ of this difference can be tested on the basis of the Chi
2 -distribution.
To account for overdispersion, we may calculate Pearson Chi
2 from observed and predicted count values
and the number of the degrees of freedom (d.f.). With Poisson-distributed data, the ratio Chi
2 /d.f. should
approximately equal unity. Values of this ratio greater than 1 indicate overdispersion. This ratio corresponds to V/m, which may be called dispersion factor (see also Eq. (2)). It can be verified relatively easily (on the basis of the definitions of Pearson Chi
2 and of the sample
variance) that the ratio Chi 2 /d.f. can be understood as
average dispersion factor over all plates and dose levels. The software STATISTICA can automatically calculate this ratio as an overdispersion parameter and use it for the correction of the variance of the estimated regression parameters and of the log-likelihood (StatSoft, 2000). The likelihood ratio Chi
2 is thereby divided by the
average dispersion factor and thus becomes smaller than without accounting for the overdispersion; so the test is then less sensitive. In summary: the estimation of the regression parameters as such is not influenced by the dispersion factor, but for statistical testing (significant contribution to the regression) the hyper-Poisson variability of the data has to be taken into account, which can be done in the described manner.
Returning to the identification of the linear region: even relatively small deviations from a strict linearity appeared to produce a significant contribution of the quadratic term in the regression function. Even with the less sensitive procedure accounting for overdispersion,
ARTICLE IN PRESS M. Roller, M. Aufderheide / Experimental and Toxicologic Pathology 60 (2008) 213–224 219
this method seems to lead to the elimination of a relatively large number of dose levels. We therefore recommend setting the level of significance to 1% instead of 5%. Fig. 3 shows dose–response relations with those dose levels that remained as the ‘‘linear region’’ after the described procedure on the 1% significance level. After visual inspection, one might easily accept the positions of the data points in terms of a linear dose–response relationship; at the same time, a sufficiently large number of data points remain in the analysis, so that a meaningful further evaluation is possible.
In each single plot of Fig. 3, the data and the regression line for the cigarette type K2R4F have been drawn in for comparison with the other cigarettes. As described above, the slopes of the straight lines in Fig. 3 have been estimated by Poisson regression; moreover, each slope value has then been related to the reference type K2R4F and expressed as a percentage of the slope value of the reference type. Thus it can be concluded, for example, that the mutagenicity of the smoke of the cigarette types A and D seems to be clearly lower than
Fig. 3. Dose–response relationships for the smoke of six cigarette typ
Ames test; results of Poisson regression accounting for overdispersio
the slope of the regression lines in relation to the reference cigarette
plots with open circles.
the mutagenicity of the smoke of the reference cigarette K2R4F, amounting to only 25% and 13%, respectively, of K2R4F after our analyses.
Testing for differences in the mutagenicity index of
two cigarette types
We found in Section ‘‘Identification of the linear region’’ that the mutagenicity of the smoke of the cigarette types A and D seems to be clearly lower than the mutagenicity of the smoke of the cigarette K2R4F. In addition, Fig. 3 shows that the types B and E also tend to produce lesser mutagenic effects than K2R4F, while type C might produce stronger effects. The question is now, whether these more or less large differences are statistically significant. This question may be answered on the basis of the procedures and statistical methods considered above. We suggest a multivariate approach for this purpose using Poisson regression and accounting for overdispersion. Perform- ing several detailed comparisons, we found that this
es on the basis of the data restricted to the linear region of the
n (see text). The percentages given (e.g. 25% and 90%) refer to
type K2R4F; the data of K2R4F are drawn in each of the six
ARTICLE IN PRESS M. Roller, M. Aufderheide / Experimental and Toxicologic Pathology 60 (2008) 213–224220
procedure is equivalent to a regression analysis on the basis of the generalized Poisson distribution according to Eq. (1) (but it does not require the formulation of a specific generalized Poisson distribution). This proce- dure accounts for the circumstance that the sample variance is larger in the available Ames test data than expected from a Poisson distribution.
The following model equation can be formulated:
counts ¼ a þ b0 � typeþðb1 þ b2 � typeÞ�dose: (5)
This equation may be used for a regression analysis with a combined data set which contains both the dose and count data of the reference type and the dose and count data of the cigarette type to be tested. An additional variable named ‘‘type’’, which has the value 0 with all doses of the reference type and has the value 1 with all doses of the test type, may be introduced into this data set. If the multivariate regression analysis leads to parameter estimates of zero for both the parameters b0 and b2, then the equation reduces to:
counts ¼ a þ b0 �dose: (6)
This means that the same linear dose–response relation- ship follows for both cigarette types (the variable ‘‘type’’ has no influence). Then there is certainly no significant difference between the two cigarette types. Of course, it is completely unlikely that the parameters b0 and b2 are both estimated to be exactly zero; probably, they will have some numerical value greater or less than zero, but again it can be tested by use of the likelihood ratio test whether the parameters contribute significantly to the regression. If the likelihood ratio-Chi
2 for the parameter
b2 is greater than the tabulated value of the Chi 2 -
distribution (with 1 degree of freedom) for a significance level of 5%, then the parameter is statistically significant on the 5% level. In this case, the slope for the reference cigarette is b1, and for the test cigarette, it is b1+b2, and this difference is statistically significant. Similarly to the identification of the linear region described above, a correction of the likelihood ratio has to be applied on the basis of the average dispersion factor also with this procedure (accounting for overdispersion).
Eq. (5) may be entered and used, e.g. in the module ‘‘Nonlinear Regression’’ of STATISTICA. However, the module ‘‘Generalized Linear Model’’ has some advantages (easier to handle, automated calculation of standard errors), but this module requires a different mode of data input, because just one parameter for each variable can be estimated in the generalized linear model. For this purpose, Eq. (5) may be simply transformed to the following equivalent equation:
counts ¼ a þ b0 � typeþ b1 �doseþ b2 � typedose:
(7)
The variables ‘‘type’’ and ‘‘dose’’ are multiplied here to another variable named ‘‘typedose’’. The input data
set must thus contain altogether four variables here: the effect variable ‘‘counts’’ and the three dose variables ‘‘dose’’, ‘‘type’’ and ‘‘typedose’’. The variable typedose has the value zero (i.e. 0 times dose) for each dose of the reference cigarette and has exactly the same value as the dose value for each dose of the test cigarette (i.e. 1 times dose).
Each of the parameters of Eq. (7) can be tested by the likelihood ratio test. STATISTICA does this automati- cally. If the parameter ‘‘a’’ is statistically significant, this means that the controls assigned to the reference cigarette have mutation rates greater than zero (cer- tainly, this will be the case and will not be surprising). If the parameter b0 is statistically significant, this may mean that the control counts of the test cigarette differ from the control counts of the reference cigarette. This is actually the case for some experiments; it also agrees with the data summarized in Table 2. However, this formal statistical significance should be interpreted with caution; it should also be borne in mind that in regression analysis the estimate of b0 is influenced to some degree by the entire dose–response relations of the cigarette types. If the parameter b1 is statistically significant, this points to a mutagenic effect of the smoke of the reference cigarette (to be expected). The parameter of greatest interest here is b2: If b2 is statistically significant, then this points to differences in the mutagenic potencies of the smoke from the reference and test cigarettes.
Table 3 contains the results for the five ‘‘test cigarettes’’ in comparison to the reference K2R4F. According to this, the slopes of all five ‘‘test cigarettes’’ are significantly different from the slope of the reference cigarette K2R4F on the 5% level and also on the 1% level without exception (p-value for the variable ‘‘type- dose’’ smaller than 0.01). With four cigarettes, potency is lower than the reference; type C shows a higher potency.
Discussion
The analysis of a large study with a specially designed Ames test showed that the relative mutagenic potency of the smoke of various cigarette types ranged from 13% to 120% related to the reference type (Table 4; Fig. 3). The differences were statistically significant. The analyses and results have to be discussed from several view points. One aspect is the type of statistical analysis. The approach chosen here has to our knowledge not been described in the literature previously, though similar approaches have been proposed (Margolin et al., 1981; Bernstein et al., 1982; Kim and Margolin, 1999). It appears most appropriate to calculate muta- genic potency from Ames test data by analysing the
ARTICLE IN PRESS
Table 3. Comparison of the mutagenic potencies of the smoke of various cigarette types with the reference type K2R4F
Variable Parameter Likelihood ratio-Chi 2
p
Estimate S.E.
K2R4F and Type A (remaining slope value: 448.2�338.2 ¼ 110.0)
Constant 21.2 0.46
Dose 448.2 8.3 7061.5 0.000000
Type �0.734 0.631 450.2 0.000000
Typedose �338.2 8.5 2290.9 0.000000
K2R4F and Type B (remaining slope value: 448.2�43.7 ¼ 404.5)
Constant 21.2 0.49
Dose 448.2 8.8 8182.3 0.0000000
Type �0.423 0.692 2.3 0.13
Typedose �43.7 12.6 11.9 0.00056
K2R4F and Type C (remaining slope value: 448.2+86.9 ¼ 535.1)
Constant 21.2 0.46
Dose 448.2 8.2 8294.0 0.000000
Type –1.503 0.634 1.5 0.21
Typedose 86.9 16.7 28.0 0.000000
K2R4F and Type D (remaining slope value: 448.2–388.1 ¼ 60.1)
Constant 21.2 0.43
Dose 448.2 7.6 3982.1 0.000000
Type �2.671 0.565 1243.8 0.000000
Typedose �388.1 7.7 3925.8 0.000000
K2R4F and Type E (remaining slope value: 448.2�181.6 ¼ 266.6)
Constant 21.2 0.52
Dose 448.2 9.3 8254.3 0.000000
Type �1.035 0.723 47.1 0.000000
Typedose �181.6 10.6 325.2 0.000000
Results of multivariate Poisson regression accounting for overdispersion, linear model, data sets restricted to the linear region (see Section ‘‘Testing
for differences in the mutagenicity index of two cigarette types’’, Eq. (7)).
Table 4. Overview of possible mutagenicity indexes for cigarette smoke and statistical significance of the differences to the smoke
of the reference K2R4F
Cigarette type Mutagenic potency a
Relative potency (%)
(related to K2R4F)
p-Value for difference
to reference type (counts per cig.equiv.) ED2 (cig.equiv.)
K2R4F 448 0.05 100 Reference
Type A 110 0.19 25 0.0000
Type B 405 0.05 90 0.0006
Type C 535 0.04 119 0.0000
Type D 60 0.31 13 0.0000
Type E 267 0.08 59 0.0000
a Both values in one row of these two columns are based on the same slope estimate. ED2 indicates the dose that leads to an increase in response by
a factor of 2; it is calculated from the corresponding slope value and control counts, e.g. for K2R4F: [21.2 counts+448 counts per cig.equiv.�0.05
cig.equiv.]/[21.2 counts] ¼ 2.
M. Roller, M. Aufderheide / Experimental and Toxicologic Pathology 60 (2008) 213–224 221
relationships between some measure of dose and the number of revertant counts by means of regression analysis. It is widely accepted that for regression analysis the so-called maximum likelihood method is preferable (Margolin et al., 1981; Bernstein et al., 1982; Kim and Margolin, 1999; US EPA, 2007). For the
application of this method, it is necessary to make assumptions on the statistical distribution of the data. The probability (density) function of the distribution is the basis for the construction of the likelihood function, negative log-likelihood and ‘‘loss function’’, respec- tively, used for fitting of the regression model.
ARTICLE IN PRESS M. Roller, M. Aufderheide / Experimental and Toxicologic Pathology 60 (2008) 213–224222
Ames test data have been regarded as Poisson- distributed where the variability is somewhat larger than expected from an exact Poisson distribution. This is called hyper-Poisson variability or overdispersion (the ratio of variance and mean V/m may be called dispersion factor). We analysed the available cigarette smoke Ames test data using various distributions, which are not all described here in detail. It turned out that very similar results were obtained when using the general Poisson distribution of Consul and Jain (1973) and when using simple Poisson regression while accounting for over- dispersion as implemented in commercial software (Sections ‘‘Regression analyses under consideration of the statistical distribution of the effect variable’’ and ‘‘Testing for differences in the mutagenicity index of two cigarette types’’). We also used the normal distribution, both in the version of constant variance and in a version of different variances dependent on dose. A similar version of regression analysis for continuous data based on the normal distribution with variance allowed to be a power function of the mean is implemented in the Benchmark Dose Software of the US EPA (2007). As expected, the version with constant variance led to a much higher degree of statistical significance of differ- ences between the cigarette types; these differences have to be evaluated as inappropriate, because actually the variances of revertant counts are not constant at different dose levels. However, regression analysis using a likelihood function based on the normal distribution with variance being a power function of the mean led to results similar to the preferred approach using Poisson distributions.
Margolin et al. (1981) and Kim and Margolin (1999) have proposed the negative binomial distribution for Ames test data. This distribution is obtained as a mixture of Poisson distributions, where the counts of each plate are assumed to be Poisson-distributed. Variance (V) is given then dependent on mean (m) and on an additional parameter c by the following relation- ship: V ¼ m(1+mc). It follows from this relationship, that the ‘‘dispersion factor’’ V/m increases with increas- ing dose, because mean m increases with increasing dose. However, such an increase was not confirmed by the data. The mean count value obtained for example for some very high dose of the reference cigarette K2R4F was 219 counts. The variance for this dose was calculated as 317. This corresponds to an actual dispersion factor for this dose of 317/219 ¼ 1.45. By using the negative binomial distribution however, a c-value of 0.05 was calculated from the pooled controls leading to a dispersion factor for the K2R4F high dose of 1+219�0.05E12. There is considerable discrepancy between the actual dispersion factor of 1.45 for this dose and the dispersion factor according to the negative binomial distribution of 12. Similar discrepancies were also found for numerous other doses. It is clear that
there was overdispersion in the data, but the actual dispersion factor was mostly in the range of 1.5–4 rather than in the range of 10–12 as estimated for relatively high doses according to the negative binomial distribu- tion. In summary, the negative binomial distribution tended to overestimate overdispersion for the higher doses. Therefore, we did not use the negative binomial distribution for the final analyses.
Besides the statistical distribution of the data, the model function for the dose–response relationship has to be selected for regression analysis. It is the characteristic of several Ames test data that there is a linear region at lower doses, a maximum of count values is reached and the revertant counts are decreasing beyond a certain dose due to substance toxicity. Examples of such curves are shown in Fig. 2. In accordance with the literature (Bernstein et al., 1982), a method was selected here to define the linear region of the dose–response relation- ship in a reproducible manner (Section ‘‘Identification of the linear region’’). This has the advantage that the potency of various substances or ‘‘influences’’ (like smoke in our case) can be compared in terms of the slopes of simple linear models (Fig. 3).
It should be noted that, with as extensive data as available here, the influence of the different statistical regression techniques on the estimate of potency was relatively small. Also, the differences between most of the examined cigarette types were so clear that there is no doubt about statistical significance. Statistically significant differences compared to the 100% reference were even found for relative mutagenicity indices of 119% and 90%. With the differences to the 100% reference being so small, statistical significance might vanish behind ‘‘practical significance’’ (i.e. it is doubtful that a reduction of the mutagenicity from 100% down to 90% is of practical importance). This may also be expressed as follows: the experimental methods and sample sizes used so far in the discussed test model provide a good framework in order to determine with high confidence differences in the mutagenicity (accord- ing to the test model) between two cigarette types, e.g. around a factor of 2.
Apart from statistics, there is a fundamental aspect when deriving measures of potency from Ames test data. It is the aim of such practices to obtain a relevant measure of the adverse effects on humans. Therefore, the results of in vitro analyses should be validated. There is certainly consensus that the ‘‘gold standard’’ against which experimental data should be validated is the experience in humans. However, it is the reason for the conduct of experimental studies that often no appropriate quantitative data from humans are avail- able. It is also widely accepted that long-term animal experiments often provide the most relevant informa- tion, particularly with respect to carcinogenicity; con- siderable analogy between such experimental data and
ARTICLE IN PRESS M. Roller, M. Aufderheide / Experimental and Toxicologic Pathology 60 (2008) 213–224 223
experience in humans has been substantiated in many cases (Wilbourn et al., 1986; Roller et al., 2006).
However, things may be quite complicated in special cases, e.g. in the case of cigarette smoke. Diesel engine emissions (DE) and environmental tobacco smoke have been compared, but important differences between the aerosols were not discussed at the same time. For example, Invernizzi et al. (2004) used a 60 m
3 garage to
assess particulate matter (PM) emission from three smouldering cigarettes, lit sequentially for 30 min, and from a 2.0 L diesel engine idling for 30 min; they reported that the tobacco smoke contributed to indoor PM concentrations up to 10-fold greater than those emitted from the diesel engine – without indicating the composition of the particulate matter. Diesel engine emissions have been shown to be carcinogenic in long- term rat experiments after inhalation and after intra- tracheal instillation (review in Roller and Pott, 2006); there are also epidemiological data indicating carcino- genicity to the lung (Garshick et al., 2004; Hoffmann and Jöckel, 2006). Altogether, the data lead to the conclusion that the carcinogenicity of diesel soot in rat lungs is largely caused by the insoluble carbon core of the particles and may be influenced only to a very small degree by carcinogenic organic compounds. It is further important to note that there is no realistic possibility to test the carcinogenicity of nanoparticles at all, if the results after instillation and the corresponding results after inhalation are evaluated as irrelevant overload in rats. The findings with diesel soot and other particles show that we have to be cautious with the interpretation of mutagenicity and genotoxicity tests. Sometimes the mutagenicity of diesel soot extracts has been used as a surrogate for the in vivo carcinogenicity of the emissions (Bünger et al., 2007). Probably this is very misleading, because of the low importance of the organic fraction of diesel particles for the observed in vivo carcinogenicity. No in vitro method is known to us which up to now has measured the carcinogenicity of insoluble particles which was observed in rat lungs. Therefore, in vitro tests with diesel particles and bio-durable nanoparticles should be validated against long-term rat studies, as far as no suitable human data are available.
On the one hand, diesel soot is a striking example where mutagenic potency should not be considered equal to carcinogenic potency. Therefore, the results with cigarette smoke also have to be interpreted cautiously with regard to carcinogenicity. On the other hand, however, there are fundamental differences between DE and cigarette smoke, both in terms of the physico- chemical characteristics of the aerosols and of the availability of data. In contrast to DE, unequivocal evidence of the carcinogenicity of cigarette smoke to the human lung is available with a very large data base, while simultaneously no long-term animal inhalation studies are known that have reproduced similar clear carcino-
genicity (Coggins, 2002; Roller et al., 2006). However, the differences between mutagenicity and carcinogenicity of cigarette smoke may be smaller than for DE. The crucial point is the existence or non-existence of an insoluble carbon core. The proportion of the carbon core was determined as 50–80% of the total mass of diesel particles (UBA, 1999). In contrast, cigarette smoke does not contain insoluble particles as far as we know, there is just condensate which dissolves in the lung. Therefore, a major part of the carcinogenic effect of cigarette smoke may also be represented in a short-term mutagenicity test, especially when only various cigarette types are com- pared. These preconditions should be borne in mind when referring to the ‘‘mutagenicity indexes’’ for cigarette smoke described in the following.
Table 4 contains some versions of a mutagenicity index based on the slope of dose–response relationships in the Ames test. All calculations were done using the dose measure introduced as ‘‘cigarette equivalents’’ ( ¼ cigarettes/dilution; Section ‘‘Selection of dose vari- able’’). Table 4 contains three different possibilities of presenting a mutagenicity index: absolute slope value, ED2 and relative slope value related to a reference aerosol. These possibilities differ only in terms of the final numerical values, they are based on the same analysis and finally contain the same information. For the variant ‘‘absolute slope value’’, e.g. K2R4F has a mutagenicity index of 448, type D has an index of 60 if the units ‘‘counts per cig.equiv.’’ are used. If percentages between 0 and 100 are preferred, then it is logical to use relative slope values related to a reference aerosol. For example, type D has a mutagenicity index of 13% related to the smoke of K2R4F, if the latter was accepted as the reference aerosol. This index has advantages of ease of communication: ‘‘the smoke of the cigarette XY has 13% of the mutagenic potency of a standard cigarette (according to the Ames test)’’. The disadvantage is that a reference aerosol must be defined. Instead of a slope value, a so-called EDx value might also be used as mutagenicity index; it is usual for Ames test data to refer to something like an ED2, i.e. the dose that doubles response. The ED2 values of Table 4 were calculated using the linear regression functions and slope values given in the same table. As with LD50 values, a high ED2 value indicates low potency and a low ED2 value indicates high potency. This kind of ‘‘thinking in opposite directions’’ may be difficult. Therefore, it seems prefer- able to express the mutagenic potency relatively to some reference substance or aerosol.
References
Aufderheide M, Gressmann H. A modified Ames assay reveals
the mutagenicity of native cigarette mainstream smoke and
its gas vapour phase. Exp Toxicol Pathol 2007;58:383–92.
ARTICLE IN PRESS M. Roller, M. Aufderheide / Experimental and Toxicologic Pathology 60 (2008) 213–224224
Bernstein L, Kaldor J, McCann J, Pike MC. An empirical
approach to the statistical analysis of mutagenesis
data from the Salmonella test. Mutat Res 1982;97:
267–81.
Bünger J, Krahl J, Munack A, Ruschel Y, Schröder O,
Emmert B, et al. Strong mutagenic effects of diesel engine
emissions using vegetable oil as fuel. Arch Toxicol 2007
[March 21, Epub ahead of print].
Coggins CR. A minireview of chronic animal inhalation
studies with mainstream cigarette smoke. Inhal Toxicol
2002;14:991–1002.
Consul P, Jain G. A generalization of the Poisson distribution.
Technometrics 1973;15:791–9.
Garshick E, Laden F, Hart JE, Rosner B, Smith THJ,
Dockery DW, et al. Lung cancer in railroad workers
exposed to diesel exhaust. Environ Health Perspect 2004;
112:1539–43.
Hoffmann B, Jöckel K-H. Lung cancer risk of occupational
exposure to diesel motor emissions and coal mine dust. Ann
NY Acad Sci 2006;1076:253–65.
Invernizzi G, Ruprecht A, Mazza R, Rossetti E, Sasco A,
Nardini S, et al. Particulate matter from tobacco versus
diesel car exhaust: an educational perspective. Tob Control
2004;13:219–21.
Kim BS, Margolin BH. Statistical methods for the Ames
Salmonella assay: a review. Mutat Res 1999;436:113–22.
Margolin BH, Kaplan N, Zeiger E. Statistical analysis of the
Ames Salmonella/microsome test. Proc Natl Acad Sci USA
1981;78:3779–83.
Myers LE, Sexton NH, Southerland LI, Wolff TJ. Regression
analysis of Ames test data. Environ Mutagen 1981;3:
575–86.
Riebe-Imre M, Aufderheide M, Emura M, Straub M, Roller M,
Mohr U, et al. Comparative studies with natural and man-
made mineral fibres in vitro and in vivo. In: Davis JMG,
Jaurand M-C, editors. Cellular and molecular effects of
mineral and synthetic dusts and fibres. Berlin, Heidelberg:
Springer; 1994. p. 273–85 [NATO ASI Series. vol. H 85].
Roller M, Pott F. Lung tumour risk estimates from rat studies
with not specifically toxic granular dusts. Ann NY Acad Sci
2006;1076:266–80.
Roller M, Akkan Z, Hassauer M, Kalberlah F. Risikoextra-
polation vom Versuchstier auf den Menschen bei Kanzer-
ogenen. Schriftenreihe der Bundesanstalt für Arbeitsschutz
und Arbeitsmedizin-Forschung-Fb 1078. Bremerhaven:
Wirtschaftsverlag NW; 2006.
StatSoft. STATISTICA for Windows, Version 5.5. Tulsa, OK,
USA: StatSoft, Inc.; 2000.
UBA. Umweltbundesamt (Ed.): Durchführung eines Risiko-
vergleichs zwischen Dieselmotoremissionen und Ottomo-
toremissionen hinsichtlich ihrer kanzerogenen und nicht-
kanzerogenen Wirkungen. Berichte 2/99-Forschungsbericht
297 61 001/01. Bearb.: Mangelsdorf I (Koordination),
Aufderheide M, Boehncke A, Melber C, Rosner G,
Heinrich U, Höpfner U, Borken J, Patyk A, Pott F, Roller
M, Schneider K, Voß J-U. Report Nr. UBA FB 99-033.
Berlin: Erich Schmidt Verlag; 1999.
US EPA (US Environmental Protection Agency). BenchMark
Dose Software (BMDS). Version 1.4.1, 2007. /http:// www.epa.gov/ncea/bmds.htmS; 2007.
Wilbourn J, Haroun L, Heseltine E, Kaldor J, Partensky C,
Vainio H. Response of experimental animals to human
carcinogens: an analysis based upon the IARC Mono-
graphs programme. Carcinogenesis 1986;7:1853–63.
- Statistical analysis of in vitro data for risk assessment - Exemplified for a case of Ames test data
- Introduction
- Experimental materials and methods
- Statistical analysis
- Selection of dose variable
- Selection of response variable
- Regression analyses under consideration of the statistical distribution of the effect variable
- Statistical distribution of the effect variable
- Regression models for the complete data sets
- Identification of the linear region
- Testing for differences in the mutagenicity index of two cigarette types
- Discussion
- References