Tuesday, January 1, 2008

Descriptive Studies: Design, Conduction and Analysis

Epidemiology is concerned with the study of distribution and determinant of disease or any health related problem in the community. Epidemiological studies can be classified as

1. Descriptive studies: In descriptive studies, the pattern of disease occurrence is described in terms of time, place and person. Descriptive studies utilise information from diverse sources of data such as census, vital statistical records, hospital records as well as national figures about consumption of food, medications and other products. Descriptive studies are conducted especially to understand the disease pattern, extent of the problem and to generate research questions and hypothesis. Descriptive studies are generally less expensive and time-consuming than the analytic studies. Types of descriptive studies are

a. Case reports and case series
b. Correlational study
c. Cross-sectional study

2. Analytic studies: are conducted to identify the risk factor(s) for diseases or evaluate the interventions for control of disease(s) or other health problems. The analytic studies include:

a. Observational studies: such as, case-control and cohort studies
b. Experimental: such as, intervention trials (clinical trial and community trial)

1.1 Case report and case series:
Case reports and case series are the most frequently published articles in the medical journals. Case reports and case series usually describe the experience of a single patient (case report) or a group of patients (case series) with a similar disease. Case report documents unusual medical occurrences and can represent the first clues in the identification of new disease or adverse effects of medications (e.g., emergence of new disease like AIDS, thalidomide tragedy and use of oral contraceptive and development of venous thrombosis etc.). Case series are collections of individual case reports, which may occur, in fairly a short period of time. Such studies are important especially to identify the beginning or presence of an epidemic in the community. These types of studies where there are unusual features of disease (or patient’s history) may lead into generation of hypothesis. Although such studies provide important clues for generation of hypothesis, they cannot be tested because of lack of an appropriate comparison group. It is therefore necessary to design an appropriate analytic study to identify the possible factor.

1.2 Correlational study:
Also called ecological study. In this study design, the investigator collects or uses information from a community (or country) as whole rather than from individuals and search for associations among various factors present in the community. Main features of such study design are:

- It usually compares mortality or disease prevalence in different groups of people or community or country
- The unit of observation is the entire community (or country)
- The estimated exposure level found in that community or geographical unit is a surrogate measure for exposure of all individuals in that unit
- Linkage between individual exposure and individual suffering for a certain disease (or death) cannot be ascertained

Once information is collected, scatter diagram of morbidity (or mortality) rates against the average exposure rates in each community (or country or geographical areas) is constructed to assess possible association between exposure and outcome of interest. Other measures of association, such as correlation coefficient, coefficient of determination, are also calculated to quantify the relationship between the variables of interest and outcome. Such studies especially help to generate hypothesis between exposure and disease association. For example, to describe the pattern of morbidity of coronary heart disease (CHD) in 1960, death rates from 44 states (in US) were collected with per capita cigarette sales. It was observed that the death rates were highest in states with most cigarette sales, lowest in those with the least sales, and intermediate in the remainder. This observation contributed to the formulation of hypothesis that cigarette smoking causes fatal CHD, which has been documented subsequently in large number of analytic studies.

The advantages of such study is that it utilises the available data from different sources for analysis, thus can be conducted with minimum time and resources. However, the limitations of such studies are a) unable to link exposure with disease among individuals; b) correlational study presents average exposure of values rather than actual individual levels and c) subjected to potential confounding bias, which cannot be controlled during analysis. For example, a correlational study found association between increased pork consumption and breast cancer. Increased pork consumption may merely be a marker for a number of other factors for the increased risk of breast cancer, such as increased dietary fat, decreased vegetable intake or higher socio-economic status of people. It is not possible to separate the effects of such potential confounding factors while analysing data.

1.3 Cross-sectional study:
Also called prevalence survey. Main objective of such study is to describe the pattern of disease prevalence (not incidence) in the community at a certain point in time as well as to test association between possible risk factors with the disease for generation of hypothesis to be tested by analytical studies. Cross sectional study measures the exposure status and disease among the individuals at the same time. In many cases it is therefore not possible to determine whether the exposure preceded or resulted from the disease. For example, a cross sectional study found association between serum retinol (vitamin A) level and colonic cancer. It is very difficult to say from the cross sectional data whether low serum retinol is responsible for development of colonic cancer or it is the consequence, due to change in dietary habits. Such type of dilemma is common virtually in all the cross sectional studies.

Cross-sectional study reflects current status of health of a community. Such study is valuable both to the public health administrator (programme personnel) as well as to the epidemiologist. Cross sectional study findings are utilised by the health planers to understand the health status of the community, extent and distribution of health problems, priority setting, efficient allocation of resources and planning for intervention, while epidemiologist utilise the findings to identify the possible risk factors to design analytical studies to confirm them and provide recommendations on preventive interventions. For example, in Bangladesh Demographic and Health Survey is conducted every two years, collecting information through household interviews from a random sample of the population. It provides valuable information on population, fertility, mortality, nutritional status, child health and maternal health for effective health care planning and administration.

Cross sectional survey can also be utilised to determine the prevalence of disease or other health outcome in a specific group of people, such as in certain occupation. Such survey provides information on occupational exposure and frequency of disease.

1.3.1 Conduction:
Cross sectional survey is usually conducted by taking a sample from the defined population of interest. If the population size is small whole population may be studied provided there is enough resources available for this. To have a representative sample of the population, sample may be selected through either of the methods such as a) simple random sampling; b) systematic sampling; c) stratified random sampling and d) cluster sampling. Once the sample is selected, data are collected from the study subjects through a) questionnaire interview; b) observation; c) physical examination; and d) lab examinations.

1.3.2 Data analysis:
As mentioned earlier, from the cross sectional study we can only calculate the prevalence of disease in the community. Data are analysed mostly in terms of descriptive statistics and presented in the form of table, graphs or charts. We can also find association between the factor(s) of interest and the disease using appropriate statistical methods.

Example:
To determine the prevalence of sputum positive tuberculosis (TB) in a community, a random sample of 400 individuals (age more than 10 years) have been selected who gave history of cough for more than 2 weeks. Morning sputum from all of them was collected and was checked for AFB (acid-fast bacillus). Out of 400 sputum collected, AFB was found in the sputum of 32 individuals. Therefore, prevalence of sputum positive TB in the community is

Prevalence of sputum +ve TB = (32  400) X 100 or 8%

Data of the survey are cross classified by sex and is given in the following table.

AFB in sputum
+Ve -Ve Total

11
(a)
239
(b) 250
(a+b)
Male


21
(c)
129
(d) 150
(b+d)
Female

Total 32 368 400

Now we can compute prevalence rates among males and females as follows:

Prevalence rate among males
= [a  (a + b)] X 100,
or, (11  250) X 100 or 4.4%;
Prevalence rate among females = [c  (c + d)] X 100,
or (21  150) X 100 or 14.0%

Data clearly indicate that the prevalence of TB is much higher among the females compared to the males. As these (prevalence rates) are not the direct measures of the disease frequency (incidence rates) among males and females, we cannot really compute the relative risk. However, we can find association between TB and sex by using the Chi-square test, formula for which is as follows:

2 = [n (ad – bc)2]  [(a + b) (c + d) (a + c) (b + d)] 

In our example, 2 = 11.7, which is much higher than the tabulated value (3.841 with df 1). Therefore, we can say that there is association between sex and sputum positive TB in the community.

In the same manner, association between other factor(s) of interest and TB can be determined.

1.3.3 Advantages and disadvantages of cross-sectional studies:

Advantages:

- Can be conducted quickly and easily with limited resources.
- Provides possible means to find association between possible risk factors and disease without question of temporality, particularly for the permanent characteristics of the individuals (e.g., sex, race, blood group etc.).

Disadvantages:

- Only provides information about disease prevalence but not the incidence of disease in the community
- It provides “snapshot” information about the community at a specific point in time
- In most instances, causal relationship between a factor and disease cannot be determined
- It requires well planned sampling scheme to generalise the findings
- There are also problem of non-response for such kind of study design

3. Hypothesis formulation from descriptive studies:

For any public health problem, first step in the search for possible solutions is to formulate a reasonable and testable hypothesis. There are three methods of hypothesis formulation about disease aetiology as described below.

a. Method of difference: involves reasoning that disease frequency is different in two sets of circumstances. If frequency of a disease is remarkably different under two different circumstances and some factor(s) can be identified in one circumstance, but is absent in the other, either this factor or its absence may have caused the disease. For example, lung cancer is very common among smokers, while it is less common among the non-smokers, leading to the hypothesis that smoking may be a risk factor for lung cancer.

Group A + Smoking → High incidence of disease
Group B without smoking → Low incidence of disease

Thus, it is hypothesised that smoking is associated with lung cancer.

b. Method of agreement: if a factor is common to a number of different circumstances, each of which have the diseases, then this factor may have some sort of relationship with the disease. For example, if cases with hepatitis A give the history of having drinking water from the same sources, before onset of illness, a hypothesis associating drinking water and hepatitis could be proposed. Other example is HIV infection, which was found to be common among injecting drug users (IDUs), haemophiliacs and recipients of blood transfusion. This finding raised the possibility of spread of HIV infection through blood and blood products.

Situation 1: A + B + C + D → Disease present
Situation 2: A + X + Y + Z → Disease present

c. Method of concomitant variation: this method involves a factor, variation (frequency and intensity) of which causes variation in frequency of disease. Correlational studies provide useful data for the formulation of such hypothesis. For example,

Group A + high intensity of sun light → High incidence of skin cancer
Group B + low intensity of sun light → Low incidence of skin cancer

No comments: