can you write 5 reflective essays(2 pages/page; arial size 12) about the topics that will be uploaded here?

profileswetiepie4611
theory-development-summer-2010_new.ppt

PROFESSOR ROBERTO N. PADUA

THEORY CONSTRUCTION AND DEVELOPMENT

COURSE OUTLINE

I. Theory,Philosophical Bases and Logic

II. Deductive Methods of Theory Development

III. Inductive Methods of Theory Development

IV. Theory Development Versus Theory Verification

Course Requirements: Workshop Outputs

LECTURE I: Theory and Philosophical Bases

1. SCIENTIFIC RESEARCH: is systematic, controlled, empirical, and critical investigation of hypothetical propositions about the presumed relationships among phenomena.


2. THEORY: is a set of interrelated constructs (concepts), definitions, and propositions that presents a systematic view of phenomena by specifying relations among variables, with the purpose of explaining, predicting, and controlling the phenomena.

DEFINITIONS

A Theory is a statement that explains why things happen as they do. There are three forms of a theory:

1. The "set-of-laws" form defines theory as a set of well-supported empirical generalizations, or "laws." Here, theory is thought of as "things we feel very certain about." This is the inductive form.

2. The "axiomatic" form defines theory as a set of interrelated propositions and definitions derived from axioms (i.e., things we feel certain about). This is the deductive form of a theory.

3. The "causal" form defines theory as a set of descriptions of causal processes. Here, theory "tells us how things work."

FUNCTIONS OF THEORY

a. EXPLANATION: provides an answer to the question "why is the fact what it is?" that is intellectually satisfying. Formal explanation: subsuming a proposition under a broader proposition which needs no explanation. It consists of a universal generalization that is assumed to be true, a particular set of circumstances, and a conclusion which asserts that an event had to occur because it was deducible from the logic of the propositions of the theory. Such explanations are deterministic/causal/nomic. Law: (x) <If Px then Qx>; Antecedent Condition: Px; Conclusion: Qx.

FUNCTIONS OF THEORY:

b. PREDICTION: proposing the occurrence of a future event given some awareness of a past or present relationship which may or may not be understood (e.g., astronomy). One can predict without explanation, but the reverse is not true. Thus explanation, rather than prediction, is the end of science.

FUNCTIONS OF THEORY

c. CONTROL: ability to intervene in a particular case or to alter the case of a particular relationship. In the pure case it implies complete understanding of elements and their relationships as well as a closed system. Less purely, it implies knowledge of the principles along which the phenomena vary.

CHARACTERISTICS OF A THEORY

ABSTRACTNESS

Abstract concepts are independent of a specific time and place. Because scientific statements must predict future events, they cannot be specific to past events. Scientists prefer theories that are as general as possible to time and place.

Abstract concepts are independent of specific circumstances or conditions. This independence permits efficiency in understanding and predicting future events. Thus, the statement, "the greater the human capital investment, the greater the life chances," contains two abstract concepts: human capital investment and life chances.

This statement can be used to derive and test a large number of related hypotheses, such as:

H1: The greater the formal education, the greater the income.

H2: The greater the job experience, the greater the likelihood of promotion.

H3: The greater the communication skills, the greater the job performance.

... and so on.

The process of science is one of moving continuously from one level of abstraction to another. Scientists "borrow" abstract statements from theories to derive hypotheses suitable to their specific study. They test these hypotheses through observation. They "return" the results of their studies to the theory by reporting to the community of scholars the efficacy of the theory in explaining their observations. Supported hypotheses provide further support for and confidence in the theory. Rejected hypotheses prompt consideration of revising the theory or noting that it is less broadly applicable than originally believed. A scientific body of knowledge is accumulated by this ongoing process of borrowing, testing, revising, and building new theories.

RELEVANCE

Empirical relevance refers to meeting two conditions of observation:

Scientific theories must be falsifiable. The distinguishing feature of science, in contrast with other epistemologies, is that its statements can, in principle, be rejected through observation.

2. Scientific theories must be supported by observations. When theories receive strong empirical support, then we gain confidence in them, which allows us to build safe bridges, send satellites into orbit, design effective crime prevention programs, etc.

THE IDEA OF A THEORY

Theories are stories, stories about how reality works. They differ from other stories in the ways described above: they are abstract, causal, and falsifiable. Nevertheless, they are stories about reality and they come from somewhere. Much has been written in the philosophy of science about induction and deduction, the twin processes by which new theories are crafted, where induction refers to designing theories by combining and raising to an abstract level empirical generalizations and deduction refers to the "great thought" about how something works.

We will rely upon Reynolds' text to describe each feature of a theory in more detail.

Theory: A set of abstract statements about reality. These statements about a phenomenon are interrelated constructs building greater understanding of the phenomenon.e.g. What would improve a nation’s quallity of life?

Proposition: One abstract statement within a theory.
Example: "The greater the human capital investment, the greater the life chances.“

Hypothesis: A specific case of the proposition.
Example: "The greater the formal education, the greater the income.’

Operational Definition: The description of how each concept will be measured.

  • The greater the years of formal schooling, the greater the total household income before taxes in 2007."
Statement Independent Variable Dependent Variable
Proposition (Abstract) Human capital investment Life chances
Hypothesis (Concrete) Formal education Income
Operational Definition Years of formal schooling Total household income before taxes in 2007

The results of the statistical test of the research hypothesis (presuming it is measured quantitatively) might lead the researcher to reject the null form of the hypothesis (i.e., "There is no relationship between formal education and income."). If so, then the results of observation lend support for the hypothesis, the proposition, the theory, and the paradigm. If the null hypothesis is not rejected, then the community of scholars will explore reasons why it was not supported, including the notion that the theory (and perhaps the paradigm) might not be a correct depiction of reality.

Concepts Concepts, the building blocks of theories, are symbols designed to convey a specific meaning to the community of scholars. They must be defined, operationalized, and reviewed by the community of scholars for meaning and accuracy. The concept self-esteem, for example, is defined as, "an individual's sense of his or her value or worth," and most often is measured using Rosenberg's Self Esteem Scale, which is widely accepted by the community of scholars.

1. Concepts are defined with either primitive or derived terms. Primitive terms cannot be defined with other symbols or language (e.g., colors, sounds, attitudes, some relationships between individuals), but can only be further described through the use of examples. A derived term is a set of primitive words and symbols that further describes a concept.

2. An abstract concept refers to two or more events (e.g., temperature, human capital investment). A concrete concept refers to a specific event (e.g., temperature of the sun, years of formal education).

3. Concepts can be measured either quantitatively or qualitatively. There is no epistemological reason to suspect that either type of measurement is more or less scientific, objective, or valid.

4. Concepts can be measured at the nominal level, indicating no inherent ranking (e.g., male, female; Christian, Hindu, Muslim, Jewish), the ordinal level, indicating ranking without a continuous ordering (e.g., large, medium, small), the interval level, indicating ranking with a continuous ordering, with no known zero-state (e.g, attitudes about same-sex marriage expressed on a 1-7 response scale), or the ratio level, indicating continuous ordered ranking with a known zero point (e.g., age in years).

1. Associational statements state a relationship without implying cause. For example, we might state that, "locus-of-control and self-esteem (two concepts with similar meanings) are related," meaning they will vary together but not necessarily cause one another.

2. Causal statements imply that x causes y (e.g., the greater the formal education, the greater the income).

3. Theoretical propositions state relationships in an abstract form (e.g., the greater the human capital investment, the greater the life chances).

4. Hypotheses state relationships in a concrete form (e.g, the greater the formal education, the greater the income).

Forms of Theory Theories can be expressed as a set of laws, in axiomatic form, or as a set of causal statements.

1. The set-of-laws format expresses relationships as a set of highly supported laws (i.e., typically in causal form). Consider, for example, the Theory of Reasoned Action, proposed by Martin Fishbein and Izak Ajzen. Within this theory we might state as one law, "the greater the attitude about the behavior, the greater the intention to engage in the behavior." All the other paths implied by the diagram would be listed as laws within the set of laws that define the theory of reasoned action.

2. The axiomatic format expresses relationships as a set of axioms. For example, within the theory of reasoned action, we might state as one axiom, "If attitude toward the behavior, then intention toward the behavior." All the other paths implied by the diagram would be listed as axioms within this format.

3. The diagram shown for the theory of reasoned action represents the causal statement form. Each diagrammed path represents a theoretical proposition. For example, we might infer from the diagram of the Theory of Reasoned Action that, "the greater the attitude about the behavior, the greater the intention to engage in the behavior."

Note Regarding the Format of Theory The typical format used in sociology to express a theory is the set of causal statements, often shown in a concise manner by the use of a diagram. In the 1980's, as part of an effort to make sociology "more scientific," sociologists began to present their theories in axiomatic format (see volumes of The American Sociological Review for examples of this effort). Sociologists learned quickly that the formatting of a theory provided few advantages toward accumulating a scientific body of knowledge; what mattered was the quality of the theory, not its formatting. Note, however, that some sociologists will argue that "theory" should be expressed either as a set of laws or in axiomatic format (see: Formal Theory in Sociology: Opportunity or Pitfall?, edited by Jerald Hage).

PHILOSOPHICAL BASES

a. EPISTEMOLOGY: How do we know what we claim to know?
1) To what extent can knowledge exist before experience?
2) To what extent is knowledge universal?
3) By what process does knowledge arise?
a) Rationalism: knowledge arises out of the sheer power of the human mind. (PLATO, ERDOS)
b) Empiricism: knowledge arises in perception (JOHN STUART MILL).
c) Constructivism: people create knowledge to function in life( LAKATOS, IMRE,).
4) Is knowledge best conceived in parts or wholes? (GESTALT)
5) To what extent is knowledge explicit?

b. ONTOLOGY: What is the nature of the phenomena we seek to know?
1) To what extent do humans make real choices?
a) Determinists (motion theory): humans are basically reactive and passive; behavior is determined by and responsive to past pressures.
b) Teleologists (action theory): people plan their behavior to meet goals; individuals create meanings, they have intentions, they make real choices.
2) To what extent are humans best understood in terms of states versus traits?
3) To what extent is human experience basically individual versus social

c. AXIOLOGY: What is the role of values in inquiry (value-conscious versus value-neutral scholarship).
1) Can theory be value-free?
2) To what extent does inquiry influence what is studied?
3) To what extent should scholarship attempt to achieve social change?

DURHEIM-QUINE PRINCIPLE

  • There are infinitely many theories that could explain the same set of data or observations.
  • Example: We all observe that the sun rises in the East and sets in the West. But:
  • A. Some theorized that the earth is flat while others theorized otherwise;
  • B. Some theorized that the earth is the center of the universe, while others believed that the sun is the center of the universe.

WORKSHOP 1:

The following exercises aim to gauge how much of the basic concepts in this lecture you have absorbed and learned. There are three(3) activities in this workshop.

1. State a theory in a field of study that you are interested in or has knowledge of.

1.1 State at least two(2) propositions based on this theory.

1.2 State at least one(1) hypothesis for every proposition that you have stated.

WORKSHOP 1 (continued)

2.. Analyze the following situation and then come up with a theory (set of at least three propositions), propositions, and hypotheses.

“ So much had been said about the deteriorating quality of Philippine higher education. In the 1960’s, the Philippines was considered Asia’s best destination for higher education and advanced studies. Thus, Philippine universities and colleges trained the technocrats of Thailand, Japan, Indonesia, Malaysia, Australia and other countries. In a recent survey of the state of education, however, the country was ranked 42nd out of 45 countries in a test for science and mathematics. Policy makers blamed the situation to many factors: short-basic education cycle, economic problems of the Philippines, non-specialized curricula in colleges and universities, brain- drain, inequitable distribution of resources etc. In the 1990’s the country re-designed its educational system to address various systemic issues leading to the poor quality of education observed. Paradoxically, the situation has not changed very much since 1994 when CHED was established.”

3. Gather the results of the other groups and copy their theories.

3.1 What are the similarities and differences in the way that the other groups derived their theories?

3.2 Did you come up with essentially the same or different theories?

3.3 How would you philosophically explain the similarities or differences in the theories that you have derived?

OUTPUT PRESENTATION :1:00-2:00 P.M. OF DAY 1

1. Make a powerpoint presentation of your results.

2. Elect a group reporter to present the output.

3. Any member of the group may respond to any questions raised by the members of the class. However, in case there are no questions, the group must also prepare a set of guide questions to steer the discussions.

4. Each group is given 15 minutes to present and respond to questions.

5. You will be graded in terms of your presentation abilities and in terms of the thoroughness with which you respond to questions.

LECTURE 2: BRIEF DIGRESSION INTO LOGIC

  • Since one of the methods for theory construction that we will study is the Deductive Method, we need to strengthen our LOGIC.
  • One handy definition for Day One of an introductory course like this is that logic is the study of argument. For the purposes of logic, an argument is not a quarrel or dispute, but an example of reasoning in which one or more statements are offered as support, justification, grounds, reasons, or evidence for another statement. The statement being supported is the conclusion of the argument, and the statements that support it are the premises of the argument.

Arguments establish the truth of conclusions relative to some premises and rules of inference. Logicians do not care whether arguments succeed psychologically in changing people's minds or convincing them. The kinks and twists of actual human reasoning are studied by psychology; the effectiveness of reasoning and its variations in persuading others are studied by rhetoric; but the correctness of reasoning (the validity of the inference) is studied by logic.

  • To assess the worth of an argument, only two aspects or properties of the argument need be considered: the truth of the premises and the validity of the reasoning from them to the conclusion. Of these, logicians study only the reasoning; they leave the question of the truth of the premises to empirical scientists and private detectives.
  • An argument is valid if the truth of its premises guarantees the truth of its conclusion.
  • Note that only arguments can be valid or invalid, not statements. Similarly, only statements can be true or false, not arguments.

Truth of Statements, Validity of Reasoning

Peter Suber , Philosophy Department , Earlham College

True Premises, False Conclusion

0.

Valid

Impossible: no valid argument can have true premises and a false conclusion.

1.

Invalid

Cats are mammals. Dogs are mammals. Therefore, dogs are cats.

True Premises, True Conclusion

2.

Valid

Cats are mammals. Tigers are cats. Therefore, tigers are mammals.

3.

Invalid

Cats are mammals. Tigers are mammals. Therefore, tigers are cats.

False Premises, False Conclusion

4.

Valid

Dogs are cats. Cats are birds. Therefore, dogs are birds.

5.

Invalid

Cats are birds. Dogs are birds. Therefore, dogs are cats.

False Premises, True Conclusion

6.

Valid

Cats are birds. Birds are mammals. Therefore, cats are mammals.

7.

Invalid

Cats are birds. Tigers are birds. Therefore, tigers are cats.

The distinction between truth and validity is the fundamental distinction of formal logic. You cannot understand how logicians see things until this distinction is clear and familiar.

PROPOSITIONAL LOGIC

  • A simple statement is one that does not contain any other statement as a part. We will use the lower-case letters, p, q, r, ..., as symbols for simple statements.
  • A compound statement is one with two or more simple statements as parts or what we will call components. A component of a compound is any whole statement that is part of a larger statement; components may themselves be compounds.
  • An operator (or connective) joins simple statements into compounds, and joins compounds into larger compounds. We will use the symbols, , · , , and to designate the sentential connectives. They are called sentential connectives because they join sentences (or what we are calling statements). The symbol, ~, is the only operator that is not a connective; it affects single statements only, and does not join statements into compounds.

Simple statements

p

"p is true"

assertion

~p

"p is false"

negation

Compounds and connectives

p http://www.earlham.edu/~peters/writing/disjunct.gifq

"either p is true, or q is true, or both"

disjunction

p · q

"both p and q are true"

conjunction

p http://www.earlham.edu/~peters/writing/matimp.gifq

"if p is true, then q is true"

implication

p http://www.earlham.edu/~peters/writing/matequiv.gifq

"p and q are either both true or both false"

equivalence

Implication statements (p http://www.earlham.edu/~peters/writing/matimp.gifq) are sometimes called conditionals , and equivalence statements (p http://www.earlham.edu/~peters/writing/matequiv.gifq) are sometimes called biconditionals .

  • The truth value of a statement is its truth or falsity. All meaningful statements have truth values, whether they are simple or compound, asserted or negated. That is, p is either true or false, ~p is either true or false, p q is either true or false, and so on.

A truth table is a complete list of the possible truth values of a statement. We use "T" to mean "true", and "F" to mean "false" (though it may be clearer and quicker to use "1" and "0" respectively).

For example, p is either true or false. So its truth table has just 2 rows:

p

T

F

But the compound, p http://www.earlham.edu/~peters/writing/disjunct.gifq, has 2 components, each of which can be true or false. So there are 4 possible combinations of truth values. The disjunction of p with q will be true as a compound whenever p is true, or q is true, or both:

p

q

p http://www.earlham.edu/~peters/writing/disjunct.gifq

T

T

T

T

F

T

F

T

T

F

F

F

If a compound has n distinct simple components, then it will have 2n rows in its truth table.

The truth table columns that define the basic connectives are as follows:

p

q

~p

~q

p http://www.earlham.edu/~peters/writing/disjunct.gifq

p · q

p http://www.earlham.edu/~peters/writing/matimp.gifq

p http://www.earlham.edu/~peters/writing/matequiv.gifq

T

T

F

F

T

T

T

T

T

F

F

T

T

F

F

F

F

T

T

F

T

F

T

F

F

F

T

T

F

F

T

T

Most statements will have some combination of T's and F's in their truth table columns; they are called contingencies . Some statements will have nothing but T's; they are called tautologies . Others will have nothing but F's; they are called contradictions . Obviously these three types of propositions exhaust the possibilities for statements that have truth table columns --which means for all truth-functional statements.

  • An argument is valid if and only if its corresponding conditional is a tautology. There are other tests for validity using truth tables. The chief alternative test searches for a counterexample or invalidating row: a possible universe (substitution instance) in which all the premises are true and the conclusion is false. If there are no counterexamples, the argument is valid; if there is even one, it is invalid.
  • Two statements are consistent if and only if their conjunction is not a contradiction.
  • Two statements are logically equivalent if and only if their truth table columns are identical --if and only if the statement of their equivalence using " " is a tautology.

  • Obviously truth tables are adequate to test validity, tautology, contradiction, contingency, consistency, and equivalence. This is important because truth tables require no ingenuity or insight, just patience and the mechanical application of rules. No matter how dumb we are, truth tables correctly constructed will always give us the right answer.

WORKSHOP 2: 2:30-4:00

  • 1. There were three prisoners. One of them is going to be executed the following day but he will not know that he will be executed until the hour of execution. Prisoner A, in an attempt to know his odds of being executed, asked the warden: “Who among us will be executed tomorrow?” The warden replied: “ I cannot tell you that. However, all I can tell you is that one of the other two prisoners will NOT be executed tomorrow”.
  • Did Prisoner A’s odds of being executed improve by this information or not?

  • 2. A king comes from a family of two children. What is the chance that the other child is a girl?
  • (Caution: The answer is not 50%. BE LOGICAL)
  • 3. Any two students in a class have some form of sociological relationship with each other. How many such sociological relationships can you deduce if there are 5 students in the class?
  • 4. Five (5) friends A,B,C,D and E are stranded in an island. Friend E was murdered on a Tuesday in their camp with a blunt object which crushed his skull instantly killing him. Friend A is scheduled to look for food every Thursday, Friday and Saturday. Friend B is scheduled to look for food every Tuesday, Friday and Saturday. Friend C’s schedule for food is Monday, Wednesday and Saturday. Friend D searches for food every Tuesday, Wednesday, Friday and Saturday. A and B are male while C and D are female. Who are your most likely suspects as murderers?

  • 5. Show by using a truth table that “p or not p” is a tautology but “p and not p” is a contradiction.
  • 6. Show that the following are equivalent statements by using a truth table:
  • 6.1 “not (p or q)” is equivalent to “not p and not q”
  • 6.2 “not (p and q” is equivalent to “not p or not q”
  • 6.3 “If p implies q” is logically equivalent to “If not q implies not p”
  • 7. Show how the proposition:”If population grows geometrically, then there will be hunger in the future” is deducible by logical inference from the Malthusian principle.
  • That is, put the Malthusian principle in symbolic logic form and do some logical operations to arrive at the statement.
  • WORKSHOP PRESENTATION: 4-5 P.M.

LECTURE 3: DEDUCTIVE SYSTEMS

DEDUCTIVE SYSTEMS

DEFINITIONS

AXIOMS AND ASSUMPTIONS

LEMMA AND LOGICAL RELATIONS DERIVED

THEORIES

COROLLARIES AND CONSEQUENCES

LECTURE 3: DEDUCTIVE METHODS

  • The process of theory construction:
  • 1.      Specify the topic
  • 2.      Specify the assumptions and axioms
  • 3.      Specify the range of phenomena
  • 4.      Specify the major concepts and variables
  • 5.      Specify the propositions, hypotheses, and relationships
  • 6.      Specify the theory

SPECIFY THE TOPIC

The first step in theory verification and/or construction is to specify the research topic of interest. Existing theories and literature related to the topic should be identified and used as guidance for determining the nature and scope of the inquiry. Since knowledge is cumulative, the inherited body of information and understanding is the takeoff point for the development of more knowledge. The practice of reviewing the literature in research papers serves this purpose of identifying relevant theories and findings or the lack of both.

EXAMPLES:

  • 1. Spread of communicable diseases.
  • 2. Quality higher education.
  • 3. Digital divide in developing countries.
  • 4. Quality of life of a nation
  • 5. Care of the elderly in developing and underdeveloped nations

SPECIFY THE ASSUMPTIONS/AXIOMS

  • The second step in theory verification and/or construction is to specify the assumptions related to the research focus. Assumptions are suppositions that are not yet tested but are considered true. In general, assumptions should make sense to most people. When in doubt, researchers should test their assumptions rather than consider them true. For example, when the telephone interview is used as a data collection method, the assumption is that it can reach a representative sample of the population of interest. If this assumption is not necessarily true, as in studies of Medicaid recipients or indigent patients, then researchers need to conduct a pretest to verity whether the telephone is a proper channel to reach the study population prior to full-scale data collection.

EXAMPLES

  • TOPIC: STATE CARE OF THE ELDERLY IN DEVELOPING AND UNDERDEVELOPED NATIONS
  • DEFINITIONS:
  • 1. Developing nations are nations whose economy are in the initial stages of industrialization; Underdeveloped nations are nations whose economy are agricuture-based.
  • AXIOMS/ASSUMPTIONS
  • 1. Developing and underdeveloped nations have fair to poor social services system.
  • 2. Developing nations prioritizes economic concerns in their national budget. Moreover, economic strategies adopted by these governments aim for productive employment.
  • 3. Developing and underdeveloped nations are characterized as religious and attach strong significance to the family.
  • 4. Population growth rates of developing and underdeveloped nations are often high: between 2.5% to 3.0% per annum. Average life expectancies range from 60 to 70 years with an sd of about 6 years.

SPECIFY THE RANGE OF PHENOMENA

  • The third step in theory verification and/or construction is to specify the range of phenomena the current research and existing theories address. For example, will the research and theories apply to people of the world or only to Filipinos or only to young Filipinos? Research or theories are more useful the greater the range of phenomena they cover, although broader theories are more difficult to construct. For one thing, data have to be collected from a wider spectrum of the population.

SPECIFY MAJOR VARIABLES AND CONCEPTS

  • The fourth step in theory verification and/or construction is to specify the major concepts and variables. Concepts are mental images or perceptions (Bailey, 1994). They may be difficult to observe directly, such as equity or ethics, or they may have referents that are easily observable, such as a hospital or a clinic. A concept that has only a single, never-changing value is called a constant. A concept that has more than one measurable value is called a variable.

  • Variables may be classified as independent and/or dependent. Generally, a variable capable of effecting change in other variables is called an independent variable. A variable whose value is dependent upon one or more other variables, but which cannot itself affect the other variables, is called a dependent variable. The dependent variable is the variable we wish to explain, and the independent variable is the hypothesized explanation. In a causal relationship, the cause is an independent variable and the effect a dependent variable. For example, since smoking causes lung cancer, smoking is an independent variable and lung cancer a dependent variable.

  • Often we can recognize a variable as independent simply because it occurs before the other variable. For example, we may find a relationship between race and level of education. Race clearly comes before schooling and, therefore, must be an independent variable. Education level can in no way influence race, since race has already been determined at birth. When one variable does not clearly precede the other, it may be difficult to designate it as dependent or independent. An example is the relationship between health status and income.

WORKSHOP 2: 10 A.M. TO 2 P.M.

  • There are two exercises in this workshop.
  • 1. THEORIES FROM CHAOS
  • 1.1 Read Robert May’s paper on Chaos in Biology.
  • 1.2 Extract the most relevant concepts from this paper in your own field of study.
  • 1.3 State some of the major assumptions and axioms that you may have in relation to applying the chaos concepts in your own field.

WORKSHOP 3: 10 A.M.- 2 P.M.DAY 2

  • 2.INTERNET WORK
  • 1. Work with your group to determine a topic which you are all interested to work on.
  • 2. Search the NET for some literature reviews on the topic. You are required to have at least three(3) downloads.
  • 3. Do a thematic review analysis. Specify your major definitions, axioms and assumptions.
  • 4. Specify the dependent and independent variables.
  • WORKSHOP PRESENTATION: 2 – 3P.M

LECTURE 4: DEDUCTIVE THEORY CONSTRUCTION CONTINUED

  • The fifth step in theory verification and/or construction is to specify the propositions, hypotheses, and relationships among the variables. A proposition is a statement about one or more concepts or variables (Bailey, 1994). Just as concepts are the building blocks of propositions, propositions are the building blocks of theories. Depending upon their use in theory building, propositions have been given different names including hypotheses, empirical generalizations, constructs, axioms, postulates, and theorems.

  • A proposition that discusses a single variable is called a univariate proposition. An example is: “Forty million of the citizens in the United States do not have any type of health insurance.” It is a univariate proposition because only one variable, “have any type of health insurance,” is contained in the statement.
  • A bivariate proposition is one that relates two variables. An example is: “The lower the population density in a county. the lower the physician-to-population ratio in that county.” It is a bivariate proposition because two variables, “population density” and “physician-to-population ratio,” are contained in the statement.

  • A proposition relating more than two variables is called a multivariate proposition. An example is: “The lower the population density in a county, the lower the physician-to-population ratio and hospital-to-population ratio in that county.” It is a multivariate proposition because three variables, “population density,” “physician-to-population ratio,” and “hospital-to-population ratio,” are contained in the statement. A multivariate proposition can be written as two or more bivariate propositions

  • For example, (1) “the lower the population density in a county, the lower the physician-to-population ratio in that county” and (2) “the lower the population density in a county, the lower the hospital-to-population ratio in that county.” This would allow for one portion of the original proposition to be rejected without rejecting the other portion, based on later statistical tests.
  •  

  • When a proposition is stated in a testable form (that we can in principle prove right or wrong through research) and predicts a particular relationship between two or more variables, it is called a hypothesis. Normative statements, or those that are opinions and value judgments, are not hypotheses. For example, the statement that every person should have access to health care is a normative statement. It is a value judgment that cannot be proved right or wrong.
  •  

SPECIFY THE THEORY

  • The final step in theory verification and/or construction is to specify the theory as applied to a particular phenomenon under investigation. The theory may be a corroborated or revised existing theory or a newly constructed theory. Theory is the result of relating the various assumptions and axioms derived.The axioms and assumptions contain variables. The formal description of a theory consists of the definitions of related concepts, the assumptions used, and a set of interrelated propositions logically formed to explain the specific topic under investigation (McCain and Sega!, 1977).

WORKSHOP 5: 3-5 P.M.

  • A. BACK TO INTERNET
  • 1. Go back to the INTERNET exercises in Workshop 4.
  • 2. Relate your various assumptions and axioms to form various theories.
  • 3. Make a report of these derived theories in class.
  • B. BACK TO STATE CARE FOR THE ELDERLY
  • 1. Weave the various assumptions and axioms into a reasonable theory.
  • 2. Generate some testable propositions and hypotheses.
  • WORKSHOP PRESENTATION: 8-9 A.M. DAY 3.

LECTURE 5: DEDUCIBILITY OF STATEMENTS

  • Our goal in this lecture is to show that theories developed deductively can be proved or disproved given the finite set of axioms or assumptions made. The method which we will use follows the axiomatic proof system .
  • Example: Prove or disprove the following :
  • 1. “In developing and underdeveloped countries, state care for the elderly will be given least priority and will instead be implicitly borne by the families concerned”.
  • 2. “In developing and underdeveloped countries, state care for the elderly will, in the long run, become an insignificant problem considering that at that point in time, the population will be mainly young.”

PROOF OF STATEMENT 1:

  • We restate our axioms for reference:
  • AXIOMS/ASSUMPTIONS
  • 1. Developing and underdeveloped nations have fair to poor social services system.
  • 2. Developing nations prioritizes economic concerns in their national budget. Moreover, economic strategies adopted by these governments aim for productive employment.
  • 3. Developing and underdeveloped nations are characterized as religious and attach strong significance to the family.
  • 4. Population growth rates of developing and underdeveloped nations are often high: between 2.5% to 3.0% per annum. Average life expectancies range from 60 to 70 years with an sd of about 6 years.

FORMAL PROOF:

  • By axiom 1, developing and underdeveloped nations have poor social services including that of providing state-supported care for the elderly.
  • By axiom 2, improvement of the social services systems of these countries, in particular, improvement of state-supported care for the elderly will not be a top priority of these governments.
  • By axiom 4, among the social services that will be on the top agenda of these governments will be the education of the ever-increasing young, school-age population and then subsequently finding jobs for them by Axiom 2.
  • Thus, state-supported care for the elderly will be a least prioritized concern of the state.
  • By axiom 3, considering the close family ties and family attachment of the citizens of these states, care for the elderly (who are members of the family) will become the concern of the nuclear families. Q.E.D

PROOF OF STATEMENT 2

  • The proof of statement 2 will require knowledge of statistics.
  • If, by axiom 4, the maximum life expectancy is 70, then the median or average age of a citizen in these countries is 35.
  • If the standard deviation is 6 years, then, at any given time, 95% of the population is between 23 years old to 47 years old.
  • From these, we conclude that the elderly (aged 65 above) in these countries constitute a mere 0.10% of the population or less than 1% of the population. (For population of 86M, then the elderly will be roughly only 86,000 people.)
  • This percentage of people falling under the category of “elderly” is small and therefore, insignificant.
  • Furthermore, a 3% increase in the younger set of people reduces the percentage of elderly by approximately 95% in the long run. QED.

COROLLARY STATEMENTS

  • From the two theories, we deduce the following corollary statement:
  • Corollary 1: Putting up a private Geriatric Center is not a profitable venture in developing and underdeveloped nations.
  • Proof:
  • The statement follows from Theory 2.

WORKSHOP 5: 10:00 – 12:00

  • 1. Consider the theories that you have developed as a group. Prove each of the statements that you made using only the finite set of assumptions and axioms that you have.
  • 2. Is your set of axioms the least number of axioms needed to prove your theories? If not, can you find the least number of axioms needed? This is called a Minimal Set of Sufficient Deductive System.
  • OUTPUT PRESENTATION: 1 – 2 PM.

LECTURE 6: DISSERTATION FORMAT AND PITFALLS

  • Theories derived by deductive methods are only as strong as their foundations or bases. If the bases are certain, then the theories derived will be as certain and and as immutable.
  • The strength of the theories derived by deductive methods rests on the researcher’s ability to scan through the “known” facts about the topic. The more facts are available, the deeper becomes your theory.
  • Previous proven theories on the topic also become part of your arsenal of axioms.

  • Old thesis and dissertation formulas actually go through the motion of scanning previous results, theories and axioms. These are then gathered in a section labelled “Theoretical Framework” or “Conceptual Framework”.
  • Graduate students in the past, use these framework as basis for formulating specific problems. They stop short of actually formulating new theories by connecting the known previous results to form new theories through a deductive inference method.

SUGGESTED FORMAT FOR DEDUCTIVE STUDIES

  • Chapter 1: Overview of the Study
  • (Introduces the topic; provides a background of the study; reviews some of the previous studies on the topic; emphasizes the importance of conducting the study; provides a logical framework for reading the study)
  • Chapter 2: Review of Previous Studies and Accepted Principles
  • (Identifies the assumptions, axioms and postulates including the variables of the study)
  • Chapter 3: Theory Formulation
  • (Develops new theories based on Chapter 2 using a deductive system; states testable hypotheses and propositions)
  • Chapter 4: Theory Verification and Validation
  • (Provides a design for testing the hypotheses; states how to tests the hypotheses; provides results of the verification through empirical data)
  • Chapter 5: Summary and Conclusion

WORKSHOP 6: 3:00-5:00 P.M.

  • Provide a detailed dissertation outline for the topic you have chosen. Replace the “generic” titles of the chapters given in the lecture with titles that will most suit your present topic.
  • OUTPUT PRESENTATION: 8:00-9:00 A.M. DAY 4

LECTURE 7: INDUCTIVE METHODS FOR THEORY DEVELOPMENT

  • Sometimes theories are created based on observations rather than on deduction from existing theories. These theories are referred to as Grounded Theories. Glaser (1992) and Strauss (1990) summarized the process of developing grounded theory as: (1) entering the field or proceeding with research without a hypothesis, (2) describing what one observes in the field, and (3) explaining why it happens on the basis of observation. These explanations become the theory, which is generated directly from observation

  • It is often very difficult to establish causality in social science research. One reason is due to the limitations of existing theories, which may not be sufficient to identify the proper causes. Another reason is that the identified causes cannot be properly controlled. Further, since much of the data in social sciences are gathered via the survey and interview method, we often cannot tell the temporal sequence of the factors of interest. Hence, we cannot be certain of the cause(s) and effect(s) and may have to treat the relationship as symmetrical without implying causality.

  • INDUCTIVE METHOD for theory development has the following key steps:
  • 1. Select a topic of interest in your field.
  • 2. Identify key factors and variables that may be discovered in reading through various studies on the topic.
  • 3. Gather data (either through the NET or survey) without any hypotheses in mind.
  • 4. Use one of the analytical data mining techniques to come up with a theory or theories.
  • 5. Formulate testable hypotheses.

Association Rule Discovery: Definition

Given a set of records each of which contain some number of items from a given collection;

Produce dependency rules which will predict occurrence of an item based on occurrences of other items

Rules Discovered:

{Milk} --> {Coke}

{Diaper, Milk} --> {Beer}

TID

Items

1

Bread, Coke, Milk

2

Beer, Bread

3

Beer, Coke, Diaper, Milk

4

Beer, Bread, Diaper, Milk

5

Coke, Diaper, Milk

  • Given a set of transactions, find rules that will predict the occurrence of an item based on the occurrences of other items in the transaction

TID

Items

1

Bread, Milk

2

Bread, Diaper, Beer, Eggs

3

Milk, Diaper, Beer, Coke

4

Bread, Milk, Diaper, Beer

5

Bread, Milk, Diaper, Coke

TID

Items

1

Bread, Milk

2

Bread, Diaper, Beer, Eggs

3

Milk, Diaper, Beer, Coke

4

Bread, Milk, Diaper, Beer

5

Bread, Milk, Diaper, Coke

Association Rule Mining Task

Item

Count

Bread

4

Coke

2

Milk

4

Beer

3

Diaper

4

Eggs

1

Itemset

Count

{Bread,Milk}

3

{Bread,Beer}

2

{Bread,Diaper}

3

{Milk,Beer}

2

{Milk,Diaper}

3

{Beer,Diaper}

3

Itemset

Count

{Bread,Milk,Diaper}

3

EXAMPLE: W e look at variables related to climate change and health concerns:

X1: average typhoons entering Philippine area of responsibility

X2: average annual temperature

X3: incidence of dengue cases per thousand population

X4: incidence of diabetes per thousand population

X5: incidence of tuberculosis per thousand population

Over the last twenty years. The data are shown below:

typhoon avtemp dengue diabetes TB

26 33 15 13 12

25 34 14 12 13

15 29 6 11 5

13 30 14 14 12

8 34 12 10 11

11 32 14 12 13

25 26 17 14 12

9 35 8 11 7

18 31 14 11 5

15 32 10 13 8

Let us discover some association rules here. We use MINITAB to convert and simplify our data as follows:

X1: Put a 1 if no. Of typhoons is more than 22 per year

X2: put a 1 if average temp. Exceeds 30

X3: Put a 1 if no. Of dengue cases exceeds 10

X4: put a 1 if no of diabetes cases exceeds 8

X5: put a 1 if no of TB cases exceeds 11

The converted data are shown below:

TYPH TEMP DENG DIAB TUBER ITEM

1 1 1 1 1 TYPH,TEMP,DENG,DIAB,TUBER

1 1 1 1 1 TYPH,TEMP, DENG, DIAB, TUBER

0 0 0 1 0 DIAB

0 0 1 1 1 DENG, DIAB, TUBER

0 1 1 1 0 TEMP, DENG,DIAB

0 1 1 1 1 TEMP, DENG, DIAB, TUBER

1 0 1 1 1 TYPH, DENG, DIAB, TUBER

0 1 0 1 0 TEMP, DIAB,

0 1 1 1 0 TEMP, DENG, DIAB

0 1 0 1 0 TEMP, DIAB

TID ITEM

1 TYPH,TEMP,DENG,DIAB,TUBER

2 TYPH,TEMP, DENG, DIAB, TUBER

3 DIAB

4 DENG, DIAB, TUBER

5 TEMP, DENG,DIAB

6 TEMP, DENG, DIAB, TUBER

7 TYPH, DENG, DIAB, TUBER

8 TEMP, DIAB,

9 TEMP, DENG, DIAB

10 TEMP, DIAB

SET OUR MINIMUM SUPPORT AT : MINSUP = 6 (ONE MORE THAN HALF THE NUMBER OF cases)

ITEM 1 SET SUPPORT

TYPHOON 3

TEMPERATURE 7

DENGUE 7

DIABETES 10

TB 5 (eliminate TYPHOON AND TB)

ITEM 2 SET SUPPORT

TEMP, DENG 6

TEMP, DIAB 7

DENG, DIAB 7 (ALL TWO ITEM SET ARE ACCEPTED)

ITEM 3 SET SUPPORT

TEMP,DENG, DIAB 5 (ELIMINATE THREE ITEM SET)

THEORIES:

1. DENGUE CASES INCREASE WITH OBSERVED RISE IN AVERAGE TEMPERATURE.

2. DIABETES CASES INCREASE WITH OBSERVED RISE IN AVERAGE TEMPERATURE

3. DIABETES AND DENGUE INCIDENCES ARE CO-EXISTENT HEALTH PROBLEMS.

WORKSHOP II

WORK ON YOUR DATA SETS AND USE ASSOCIATION ANALYSIS TO DEVELOP YOUR VARIOUS THEORIES.

  • THEORY:
  • “THE INTENSITY AND SPREAD OF VECTOR BORNE DISEASES VARY WITH THE INTENSITY OF CLIMATE CHANGE”
  • DO WORKSHOP 7 ON NEXT SLIDE.
  • OUTPUT PRESENTATION IS 11:00-12:00

WORKSHOP 7: DO AN ASSOCIATION ANALYSIS FOR THE FOLLOWING STRESS LEVEL DATA: 10:00-11:00 Day 4

age gender salary exercise diet parents’history stress level

52 1 37664 3 5 1 21.4888

31 1 23400 1 2 2 12.1500

29 0 4000 3 4 1 6.4177

27 1 39379 4 4 2 22.2606

60 0 7000 1 2 2 6.1745

52 0 11500 1 5 2 9.6898

45 0 39715 1 4 1 21.2718

28 1 14344 3 1 1 9.0148

42 0 7229 2 3 2 6.5931

39 1 14609 2 2 1 8.3541

52 1 23678 1 2 1 12.6951

32 1 23758 3 2 0 13.3311

38 0 35271 4 4 1 21.1320

39 0 26000 3 2 0 14.9530

48 1 24805 2 2 0 13.6223

34 1 38579 3 3 2 20.5406

46 1 18000 5 5 0 18.1015

26 0 13864 1 2 1 7.7588

42 0 17119 4 4 1 12.5436

32 1 9516 2 5 1 7.9422

LECTURE 8:ANOMALY DETECTION VIA REGRESSION DIAGNOSTICS

  • Regression analysis is concerned with determining a relationship between a dependent variable (y) and a set of independent variables x1,x2,x3,…,xp. It is a confirmatory (validation) statistical technique.
  • However, as a by-product of regression analysis, software packages provide regression diagnostics. These diagnostics signal the presence of outliers in a given data set.
  • These outliers (or anomalous observations) provide a hint for the development of theories.

EXAMPLE:

  • Mirasol (2009) conducted a study on the management of protected areas (MPA) in the Philippines. To this end, he defined the following variables:
  • 1. Mgt: refers to the extent to which an MPA had been properly managed (1=poor to 5 = excellent)
  • 2. Type: refers to the type of MPA (0 = marine, 1=terrestrial)
  • 3. NGO: refers to the existence of NGO support
  • 4. INC: refers to the income derived by communities in nearby non-MPA.

DATA

TYPE MGT NG0 INC.

1 5 0 4

1 4 0 4

1 5 0 5

1 4 0 3

1 3 0 3

1 4 0 5

1 5 1 5

1 4 1 4

1 3 0 4

1 3 1 3

0 2 0 2

0 2 0 2

0 3 1 5

0 4 1 5

0 2 0 2

0 1 0 2

0 5 1 1

0 2 0 1

0 2 0 2

0 1 0 2

REGRESSION:

  • We looked at the income (y) in relation to the other variables as independent variables.

The regression equation is

INC. = 1.57 + 1.16 TYPE + 0.272 MGT + 0.594 NG0

Predictor Coef SE Coef T P

Constant 1.5690 0.7091 2.21 0.042

TYPE 1.1648 0.5792 2.01 0.046

MGT 0.2720 0.1185 2.29 0.020

NG0 0.5939 0.6358 0.93 0.364

  • The anomalous observations are:
  • Both unusual observations are marine protected areas. A typical MPA should have poor income but observation 13 has a high income. Further perusal of observation 13 revealed that it had good management (3) and is supported by an NGO(1). This MPA is in Baliangao, Misamis Occidental.
  • Observation 17 became unusual despite its low income because it had excellent management (5) and supported by an NGO(1).Observation 17 is an MPA off the coastal area of Sulu.

Unusual Observations

Obs TYPE INC. Fit SE Fit Residual St Resid

13 0.00 5.000 2.979 0.522 2.021 2.05R

17 0.00 1.000 3.523 0.734 -2.523 -3.01R

  • The proposition:” Terrestrial protected areas are better managed and consequently more economically viable than marine protected areas” is modified when the unusual observations are analyzed.
  • The modification may be stated thus:” Marine protected areas in relatively peaceful location which are well-managed and with sufficient external NGO support can perform as well as terrestrial protected areas economically.”

WORKSHOP

  • The following is an actual dissertation data that attempted to determine the worth of accreditation as a means to improve quality of a higher education program. Data from 25 higher education institutions were obtained on: accrediting agency, level of accreditation of teacher education program, and performance in LET as surrogate to the term “quality”. The data are provided on the next slide.
  • Perform a regression analysis on level of accreditation vs. Quality.
  • Note the unusual observations and correspondingly revise your theory.
  • Note: 3 = PAASCU, 2 = PACU-COA, 1 = AACCUP

agency level LET

1 2 78

1 2 77

1 2 78

1 1 64

1 3 81

1 3 82

2 3 66

2 4 65

2 3 78

2 2 65

2 3 76

2 2 80

2 3 80

3 4 90

3 3 88

3 2 78

3 1 73

3 3 89

3 2 80

3 3 89

3 4 91

3 3 90

3 4 92

3 2 78

3 2 79

LECTURE 3: CLUSTER ANALYSIS

  • Purpose: Cluster Analysis is a multivariate exploratory data analysis method which aims to group individuals according to their similarities on certain measurable factors and variables.
  • Use:
  • 1. The main use of cluster analysis is to generate tentative generalizations and theories based on the groupings of individuals i.e. why a subset of individuals is grouped together and how this subset of individuals differ from other subsets of individuals. It is very useful in policy formulation;
  • 2. It is also used in Biological Sciences for species identification i.e. whenever a new organism is discovered, it may be useful to identify where it belongs in terms of family, genus or species.

*

PRACTICAL SITUATIONS

  • We deal with clustering in almost every aspect of daily life. For example, a group of diners sharing the same table in a restaurant may be regarded as a cluster of people. In food stores items of similar nature, such as different types of meat or vegetables are displayed in the same or nearby locations. There is a countless number of examples in which clustering plays an important role.
  • For instance, biologists have to organize the different species of animals before a meaningful description of the differences between animals is possible. According to the modern system employed in biology, man belongs to the primates, the mammals, the amniotes, the vertebrates, and the animals. Note how in this classification, the higher the level of aggregation the less similar are the members in the respective class. Man has more in common with all other primates (e.g., apes) than it does with the more "distant" members of the mammals (e.g., dogs), etc.

DATA REQUIREMENTS FOR CLUSTER ANALYSIS

  • In cluster analysis, we will need to measure a set of characteristics (variables or factors) per individual. Denote the characteristics as : x1, x2, …, xp. . For example:
  • X1 = income
  • X2 = nutritional status
  • X3 = level of sanitation of home
  • X4 = availability of water
  • For each individual, we measure four (4) characteristics.

The Concept of “Distance” Between Individuals

  • The concept of “distance” between individuals is crucial in cluster analysis. We will need to know how similar or dissimilar a certain individual A is from an individual B. The distance d(A,B) is a measure of how far individual A is from individual B i.e. if d(A,B) = 0, then A and B are essentially the same individual with respect to the measured characteristics.

A BLAST FROM THE PAST!!!! The Geometry Nightmare Revisited

  • Let us say that for two individuals A and B the characteristic measured is: x1 = body odor with 0 = no odor, and 5 = very odorous. For individual A , x 1 = 0 while for individual B, x 1 = 5, Then it is natural to estimate the “odorous distance” between A and B by:
  • d(A,B) = /0 – 5/ = /-5/ = 5,
  • i.e. the two individuals are very “far” from each other with respect to body odor!!!

Two Characteristics

  • Now, let us measure two characteristics per individual: x 1 = body odor, and X2 = color of skin, where color of skin is 1 = light color, to 5 = very dark color. Thus,
  • A = (0, 5) and B = (5, 1) means that A has no body odor (x1 =0) and has a very dark skin color (x 2 = 5) while B has very strong body odor (x 1 = 5) and has a very light skin color (x2 = 1). Then, from geometry:
  • d(A,B) = [(0-5)2 + (5-1)2]1/2 = (41)1/2 = 6.41.
  • Individuals C = (2,2) and D = (2, 3) would have:
  • d(C,D) = [(2-2)2 + (2-3)2]1/2 = (0 +1)1/2 = 1.
  • So, C and D are more similar to each other than are A and B.

CLUSTERING ALGORITHM

  • k-Means Clustering
  • General logic Suppose that you already have hypotheses concerning the number of clusters in your cases or variables. You may want to "tell" the computer to form exactly 3 clusters that are to be as distinct as possible. This is the type of research question that can be addressed by the k- means clustering algorithm. In general, the k-means method will produce exactly k different clusters of greatest possible distinction. It should be mentioned that the best number of clusters k leading to the greatest separation (distance) is not known as a priori and must be computed from the data. (Good research for the theoretical mathematician)

A WORKED EXAMPLE

  • We are looking at “QUALITY OF EDUCATION” and we are getting data from 15 universities in the Philippines. We decided to look into:
  • X1 = performance in licensure exams
  • X2 = average tuition rate per unit
  • X3 = enrollment size
  • X4 = Percentage of Ph.D.’s in faculty
  • X5 = acceptance/rejection rate per hundred
  • The data are shown on the next page.

THE DATA SET

  • Exam Tuition Enrolment Ph.D. rejection
  • 87 700 6000 60 0.80
  • 85 620 5500 50 0.75
  • 83 600 5000 45 0.77
  • 82 600 5400 48 0.74
  • 83 610 6250 50 0.80
  • 78 450 8000 36 0.65
  • 77 400 7600 32 0.54
  • 76 410 7700 37 0.45
  • 78 460 8900 32 0.50
  • 80 500 9000 30 0.48
  • 72 200 8000 8 0.20
  • 75 250 9000 10 0.10
  • 70 300 9000 7 0.15
  • 67 260 11000 10 0.15
  • 68 200 10000 5 0.05

STEPS

GO TO “STAT”

GO TO “MULTIVARIATE”

GO TO “CLUSTER OBSERVATIONS”

CLICK ALL THE VARIABLES

CLICK “SHOW DENDROGRAM”

Input “number of clusters = 3”

Click “OK”

THE DENDROGRAM

27.txt

��������Cluo 'Exam'-'rejection';���������������������������������������������������������������������������������������������������������������������/������}.��Cluo 'Exam'-'rejection';��������������������������������������������������������������������������������������������������������;; HMF V1.24 TEXT ;; (Microsoft Win32 Intel 386) HOOPS 5.00-17 I.M. 3.00-17 (Selectability "windows=off,geometry=on") (Visibility "on") (Color_By_Index "Geometry,Face Contrast" 1) (Color_By_Index "Window" 0) (Window_Frame "off") (Window -1 1 -1 1) (Camera (0 0 -5) (0 0 0) (0 1 0) 2 2 "Stretched") ;; (Driver_Options "no backing storeno borderno control areano debug,disable in ;; put,no double-bufferingno double bufferingno fixed colors,no force black-and ;; -whiteno force black and whiteno gamma correctionlight scaling=0,no locater ;; transform,no pen speed,no physical size,no subscreen creatingno subscreen mo ;; vingno subscreen resizingsubscreen stretchingno output format,no use colorma ;; p id") (Edge_Pattern "---") (Edge_Weight 1) (Face_Pattern "solid") (Heuristics "no related selection limit") (Line_Pattern "---") (Line_Weight 1) (Marker_Size 0.421875) (Marker_Symbol ".") (Text_Font "name=arial-gdi-vector,no transforms,rotation=follow path") (User_Options "mtb aspect ratio=0.675953,graphicsversion=6,worksheettitle=\"Wor ksheet 1\",optiplot=0,builtin=0,statguideid=0,toplayer=0,angle=0,arrowdir=0,arr owstyle=0,polygon=0,isdata=0,textfollowpath=1,ldfill=0,solidfill=0,3d=0,usebitm ap=0,canbrush=0,brushrows=0,light scaling=0.00000,sessionline=58") (Segment "include" ()) (Front ((Segment "figure1" ( (Window_Pattern "clear") (Window -1 1 -1 1) (User_Options "viewinfigurecoord=0") (Front ((Segment "region" ( (Front ((Segment "figure box" ( (Visibility "polygons=off,lines=off") (Color_By_Index "Face" 0) (Color_By_Index "Face Contrast,Line,Edge" 1) (Edge_Pattern "---") (Edge_Weight 1) (Face_Pattern "solid") (Line_Pattern "---") (Line_Weight 1) (User_Options "solidfill=1") (Segment "" ( (Polygon ((-0.99995 -0.99995 0) (0.99995 -0.99995 0) (0.99995 0.99995 0) (-0.99995 0.99995 0) (-0.99995 -0.99995 0))))))) (Segment "data box" ( (Visibility "faces=off") (Color_By_Index "Face" 0) (Color_By_Index "Face Contrast,Line,Edge" 1) (Edge_Pattern "---") (Edge_Weight 1) (Face_Pattern "solid") (Line_Pattern "---") (Line_Weight 1) (User_Options "solidfill=1") (Segment "" ( (Polygon ((-0.749962 -0.59997 0) (0.899955 -0.59997 0) (0.899955 0.59997 0) (-0.749962 0.59997 0) (-0.749962 -0.59997 0))))))) (Segment "legend box" ()) (Segment "legend" ( (Window_Pattern "clear") (Window -1 1 -1 1) (User_Options "viewinfigurecoord=1"))))))) (Segment "object" ( (Front ((Segment "frame" ( (Window_Pattern "clear") (Window -1 1 -1 1) (Front ((Segment "tick" ( (Front ((Segment "set1" ( (Color_By_Index "Face Contrast,Line,Text,Edge" 1) (Edge_Pattern "---") (Edge_Weight 1) (Line_Pattern "---") (Line_Weight 1) (Text_Alignment "^*") (Text_Font "name=arial-gdi-vector,size=0.03385 sru") (Segment "" ( (Text 0.734963 -0.669966 0 "15"))) (Segment "" ( (Text 0.624969 -0.669966 0 "13"))) (Segment "" ( (Text 0.514974 -0.669966 0 "12"))) (Segment "" ( (Text 0.40498 -0.669966 0 "10"))) (Segment "" ( (Text 0.294985 -0.669966 0 "9"))) (Segment "" ( (Text 0.184991 -0.669966 0 "8"))) (Segment "" ( (Text 0.0749962 -0.669966 0 "7"))) (Segment "" ( (Text -0.0349982 -0.669966 0 "11"))) (Segment "" ( (Text -0.144993 -0.669966 0 "6"))) (Segment "major" ( (Segment "" ( (Polyline ((0.734963 -0.59997 0) (0.734963 -0.639968 0) )))) (Segment "" ( (Polyline ((0.624969 -0.59997 0) (0.624969 -0.639968 0) )))) (Segment "" ( (Polyline ((0.514974 -0.59997 0) (0.514974 -0.639968 0) )))) (Segment "" ( (Polyline ((0.40498 -0.59997 0) (0.40498 -0.639968 0))) )) (Segment "" ( (Polyline ((0.294985 -0.59997 0) (0.294985 -0.639968 0) )))) (Segment "" ( (Polyline ((0.184991 -0.59997 0) (0.184991 -0.639968 0) )))) (Segment "" ( (Polyline ((0.0749962 -0.59997 0) (0.0749962 -0.639968 0))))) (Segment "" ( (Polyline ((-0.0349982 -0.59997 0) (-0.0349982 -0.639968 0))))) (Segment "" ( (Polyline ((-0.144993 -0.59997 0) (-0.144993 -0.639968 0))))))))) (Segment "set2" ( (Color_By_Index "Face Contrast,Line,Text,Edge" 1) (Edge_Pattern "---") (Edge_Weight 1) (Line_Pattern "---") (Line_Weight 1) (Text_Alignment "^*") (Text_Font "name=arial-gdi-vector,size=0.03385 sru") (Segment "" ( (Text 0.844958 -0.669966 0 "14"))) (Segment "major" ( (Segment "" ( (Polyline ((0.844958 -0.59997 0) (0.844958 -0.639968 0) )))))))) (Segment "set3" ( (Color_By_Index "Face Contrast,Line,Text,Edge" 1) (Edge_Pattern "---") (Edge_Weight 1) (Line_Pattern "---") (Line_Weight 1) (Text_Alignment "^*") (Text_Font "name=arial-gdi-vector,size=0.03385 sru") (Segment "" ( (Text -0.254987 -0.669966 0 "3"))) (Segment "" ( (Text -0.364982 -0.669966 0 "4"))) (Segment "" ( (Text -0.474976 -0.669966 0 "2"))) (Segment "" ( (Text -0.584971 -0.669966 0 "5"))) (Segment "" ( (Text -0.694965 -0.669966 0 "1"))) (Segment "major" ( (Segment "" ( (Polyline ((-0.254987 -0.59997 0) (-0.254987 -0.639968 0))))) (Segment "" ( (Polyline ((-0.364982 -0.59997 0) (-0.364982 -0.639968 0))))) (Segment "" ( (Polyline ((-0.474976 -0.59997 0) (-0.474976 -0.639968 0))))) (Segment "" ( (Polyline ((-0.584971 -0.59997 0) (-0.584971 -0.639968 0))))) (Segment "" ( (Polyline ((-0.694965 -0.59997 0) (-0.694965 -0.639968 0))))))))) (Segment "set4" ( (Color_By_Index "Face Contrast,Line,Text,Edge" 1) (Edge_Pattern "---") (Edge_Weight 1) (Line_Pattern "---") (Line_Weight 1) (Text_Alignment "*>") (Text_Font "name=arial-gdi-vector,size=0.03385 sru") (Segment "" ( (Text -0.819959 0.54283 0 " 77.26"))) (Segment "" ( (Text -0.819959 0.165706 0 " 84.84"))) (Segment "" ( (Text -0.819959 -0.222846 0 " 92.42"))) (Segment "" ( (Text -0.819959 -0.59997 0 " 100.00"))) (Segment "major" ( (Segment "" ( (Polyline ((-0.749962 0.54283 0) (-0.78996 0.54283 0))) )) (Segment "" ( (Polyline ((-0.749962 0.165706 0) (-0.78996 0.165706 0) )))) (Segment "" ( (Polyline ((-0.749962 -0.222846 0) (-0.78996 -0.222846 0))))) (Segment "" ( (Polyline ((-0.749962 -0.59997 0) (-0.78996 -0.59997 0) )))))))))))) (Segment "grid" ()) (Segment "reference" ()) (Segment "axis" ( (Front ((Segment "dend1" ( (Color_By_Index "Text" 1) (Text_Alignment "*<") (Text_Font "name=arial-gdi-vector,size=0.04232 sru") (Segment "" ( (Text -0.979951 0.699965 0 "Similarity"))) (Segment "" ( (Text -0.0905559 -0.79996 0 "Observations"))))))))))))) (Segment "data" ( (Window_Pattern "clear") (Window -1 1 -1 1) (User_Options "isdata=1,viewinfigurecoord=1") (Front ((Segment "dend1" ( (Color_By_Index "Face Contrast,Line,Edge" 1) (Edge_Pattern "---") (Edge_Weight 1) (Line_Pattern "---") (Line_Weight 1) (Segment "" ( (Polyline ((0.734963 0.237484 0) (0.734963 -0.59997 0))))) (Segment "" ( (Polyline ((0.239988 0.237484 0) (0.239988 0.152824 0))))) (Segment "" ( (Polyline ((0.239988 0.237484 0) (0.734963 0.237484 0))))) (Segment "" ( (Polyline ((0.459977 0.152824 0) (0.459977 -0.440642 0))))) (Segment "" ( (Polyline ((0.019999 0.152824 0) (0.019999 -0.346829 0))))) (Segment "" ( (Polyline ((0.019999 0.152824 0) (0.459977 0.152824 0))))) (Segment "" ( (Polyline ((0.129994 -0.346829 0) (0.129994 -0.515807 0))))) (Segment "" ( (Polyline ((-0.0899955 -0.346829 0) (-0.0899955 -0.389508 0)) ))) (Segment "" ( (Polyline ((-0.0899955 -0.346829 0) (0.129994 -0.346829 0)))) ) (Segment "" ( (Polyline ((-0.0349982 -0.389508 0) (-0.0349982 -0.59997 0))) )) (Segment "" ( (Polyline ((-0.144993 -0.389508 0) (-0.144993 -0.59997 0))))) (Segment "" ( (Polyline ((-0.144993 -0.389508 0) (-0.0349982 -0.389508 0))) )) (Segment "" ( (Polyline ((0.569971 -0.440642 0) (0.569971 -0.557868 0))))) (Segment "" ( (Polyline ((0.349982 -0.440642 0) (0.349982 -0.509858 0))))) (Segment "" ( (Polyline ((0.349982 -0.440642 0) (0.569971 -0.440642 0))))) (Segment "" ( (Polyline ((0.40498 -0.509858 0) (0.40498 -0.59997 0))))) (Segment "" ( (Polyline ((0.294985 -0.509858 0) (0.294985 -0.59997 0))))) (Segment "" ( (Polyline ((0.294985 -0.509858 0) (0.40498 -0.509858 0))))) (Segment "" ( (Polyline ((0.184991 -0.515807 0) (0.184991 -0.59997 0))))) (Segment "" ( (Polyline ((0.0749962 -0.515807 0) (0.0749962 -0.59997 0))))) (Segment "" ( (Polyline ((0.0749962 -0.515807 0) (0.184991 -0.515807 0))))) (Segment "" ( (Polyline ((0.624969 -0.557868 0) (0.624969 -0.59997 0))))) (Segment "" ( (Polyline ((0.514974 -0.557868 0) (0.514974 -0.59997 0))))) (Segment "" ( (Polyline ((0.514974 -0.557868 0) (0.624969 -0.557868 0)))))) ) (Segment "dend2" ( (Color_By_Index "Face Contrast,Line,Edge" 1) (Edge_Pattern "---") (Edge_Weight 1) (Line_Pattern "---") (Line_Weight 1) (Segment "" ( (Polyline ((0.844958 0.237923 0) (0.844958 -0.59997 0))))) (Segment "" ( (Polyline ((0.844958 0.237923 0) (0.844958 -0.59997 0))))) (Segment "" ( (Polyline ((0.844958 0.237923 0) (0.844958 0.237923 0))))))) (Segment "dend3" ( (Color_By_Index "Face Contrast,Line,Edge" 1) (Edge_Pattern "---") (Edge_Weight 1) (Line_Pattern "---") (Line_Weight 1) (Segment "" ( (Polyline ((-0.337483 -0.176376 0) (-0.337483 -0.265408 0)))) ) (Segment "" ( (Polyline ((-0.639968 -0.176376 0) (-0.639968 -0.377556 0)))) ) (Segment "" ( (Polyline ((-0.639968 -0.176376 0) (-0.337483 -0.176376 0)))) ) (Segment "" ( (Polyline ((-0.254987 -0.265408 0) (-0.254987 -0.59997 0))))) (Segment "" ( (Polyline ((-0.419979 -0.265408 0) (-0.419979 -0.514622 0)))) ) (Segment "" ( (Polyline ((-0.419979 -0.265408 0) (-0.254987 -0.265408 0)))) ) (Segment "" ( (Polyline ((-0.584971 -0.377556 0) (-0.584971 -0.59997 0))))) (Segment "" ( (Polyline ((-0.694965 -0.377556 0) (-0.694965 -0.59997 0))))) (Segment "" ( (Polyline ((-0.694965 -0.377556 0) (-0.584971 -0.377556 0)))) ) (Segment "" ( (Polyline ((-0.364982 -0.514622 0) (-0.364982 -0.59997 0))))) (Segment "" ( (Polyline ((-0.474976 -0.514622 0) (-0.474976 -0.59997 0))))) (Segment "" ( (Polyline ((-0.474976 -0.514622 0) (-0.364982 -0.514622 0)))) ))) (Segment "dend4" ( (Color_By_Index "Face Contrast,Line,Edge" 1) (Edge_Pattern "---") (Edge_Weight 1) (Line_Pattern "---") (Line_Weight 1) (Segment "" ( (Polyline ((0.666217 0.54283 0) (0.666217 0.237923 0))))) (Segment "" ( (Polyline ((-0.488726 0.54283 0) (-0.488726 -0.176376 0))))) (Segment "" ( (Polyline ((-0.488726 0.54283 0) (0.666217 0.54283 0))))) (Segment "" ( (Polyline ((0.487476 0.237923 0) (0.487476 0.237484 0))))) (Segment "" ( (Polyline ((0.487476 0.237923 0) (0.844958 0.237923 0)))))))) ))))))) (Segment "labels" ( (Window_Pattern "clear") (Window -1 1 -1 1))) (Segment "annotation" ( (Window_Pattern "clear") (Window -1 1 -1 1))))))) (Segment "annotation" ( (Window_Pattern "clear") (Window -1 1 -1 1) (User_Options "toplayer=1")))))

OUTPUT 1

  • Final Partition
  • Number of clusters: 3
  • Number of Within cluster Average distance Maximum distance
  • observations sum of squares from centroid from centroid
  • Cluster1 5 995263.203 397.972 630.562
  • Cluster2 9 5165381.279 680.087 1430.462
  • Cluster3 1 0.000 0.000 0.000

OUTPUT 2

  • Cluster Centroids
  • Variable Cluster1 Cluster2 Cluster3 Grand centrd
  • Exam 84.0000 74.8889 67.0000 77.4000
  • Tuition 626.0000 352.2222 260.0000 437.3333
  • Enrolment 5630.0000 8577.7778 11000.00 7756.6667
  • Ph.D. 50.6000 21.8889 10.0000 30.6667
  • rejection 0.7720 0.3467 0.1500 0.4753
  • Distances Between Cluster Centroids
  • Cluster1 Cluster2 Cluster3
  • Cluster1 0.0000 2960.6174 5382.6382
  • Cluster2 2960.6174 0.0000 2424.0192
  • Cluster3 5382.6382 2424.0192 0.0000

HYPOTHESES AND PROPOSITIONS

  • 1. Institutions with a good number of faculty with advanced degrees have better quality of education.
  • 2. Institutions with higher tuition also have lower enrollment.
  • 3. Institutions with lower enrollment have lower student to faculty ratio and turn out to have better quality also.

workshop

  • PART A
  • 1. Add one more proposition or hypothesis in the last slide of this lecture. Produce a working theory for the entire exercise.
  • 2. Perform a regression analysis for all the hypothesized relationships between quality and the other factors or variables, singly and in combination.
  • 3. If there are any anomalous observations, reformulate the theory that you have made in Problem 1.

Workshop

  • PART B.
  • Human Development Index (HDI) is a composite index measuring average achievement in three basic dimensions of human development: a long and healthy life, knowledge, and decent standard of living. In 2003, a survey was conducted to determine the HDI of several countries (Fukuda-Parr, 2003). In this survey, the Philippines was ranked 85th out of 175 countries in the world. Data for fifteen (15) selected countries are shown on the next page. Perform a cluster analysis on this data set and formulate tentative theories about Human Development Indices of countries worldwide.

HUMAN DEVELOPMENT INDEX DATA (Fukuda-Parr,2003, Milenium Development Goals)

Country HDI MALNUTRITION LITERACY POVERTY POLITICAL
1. Norway 0.994 0.01 0.99 0.02 0.99
2. Japan 0.933 0.01 0.94 0.04 0.94
3. Germany 0.921 0.03 0.96 0.05 0.92
4. Singapore 0.884 0.04 0.87 0.03 0.95
5. Brunei 0.872 0.07 0.89 0.02 0.99
6. Malaysia 0.79 0.08 0.83 0.06 0.92
7. Thailand 0.768 0.08 0.88 0.05 0.88
8. Philippines 0.751 0.00 0.9 0.1 0.84
9. Vietnam 0.688 0.08 0.83 0.11 0.88
10. Indonesia 0.682 0.09 0.80 0.11 0.85
11. Cambodia 0.556 0.12 0.64 0.15 0.84
12. Myanmar 0.549 0.20 0.72 0.22 0.8
13. Sierra Leone 0.275 0.25 0.41 0.25 0.77
14. USA 0.94 0.01 0.97 0.04 0.98
15. Ethiopia 0.33 0.21 0.45 0.21 0.80

  • C. Read the paper on the Normal Distribution. Provide an in-depth critique of the arguments used in the paper. Take careful note of the following:
  • 1. Is the present use of the normal distribution in education justified on the basis of the historical background provided by the paper?
  • 2. What is the main “thesis” of the paper? Is the author’s argument convincing and free of contradictions?
  • 3. If you were to rewrite the paper, how would you modify the author’s work?

Concepts

Concepts, the building blocks of theories, are symbols designed to convey a specific

meaning to the community of scholars. They must be defined, operationalized, and

reviewed by the community of scholars for meaning and accuracy. The concept self-esteem,

for example, is defined as, "an individual's sense of his or her value or worth," and most

often is measured using Rosenberg's Self Esteem Scale , which is widely accepted by the

community of scholars.

1. Concepts are defined with either primitive or derived terms. Primitive terms cannot be

defined with other symbols or language (e.g., colors, sounds, attitudes, some

relationships between individuals ), but can only be further described through the use of

examples. A derived term is a set of primitive words and symbols that further describes

a concept.

2. An abstract concept refers to two or more events (e.g., temperature, human capital

investment). A concrete concept refers to a specific event (e.g., temperature of the

sun, years of formal education).

3. Concepts can be measured either quantitatively or qualitatively. There is no

epistemological reason to suspect that either type of measurement is more or less

scientific, objective, or valid.

4. Concepts can be measured at the nominal level, indicating no inherent ranking (e.g.,

male, female; Christian, Hindu, Muslim, Jewish), the ordinal level, indicating ranking

without a continuous ordering (e.g., large, m edium, small), the interval level,

indicating ranking with a continuous ordering, with no known zero -state (e.g,

attitudes about same-sex marriage expressed on a 1 -7 response scale), or the ratio

level, indicating continuous ordered ranking with a known ze ro point (e.g., age in

years).

1. Associational statements state a relationship without implying cause. For example,

we might state that, "locus -of-control and self-esteem (two concepts with similar

meanings) are related," meaning they will vary together but not necessarily cause

one another.

2. Causal statements imply that x causes y (e.g., the greater the formal education, the

greater the income).

3. Theoretical propositions state relationships in an abstract form (e.g., the greater the

human capital investment, the greater the life chances).

4. Hypotheses state relationships in a concrete form (e.g, the greater the formal

education, the greater the income).

Forms of Theory

Theories can be expressed as a set of laws, in axiomatic form, or as a set of causal

statements.

1. The set-of-laws format expresses relationships as a set of highly supported laws (i.e.,

typically in causal form). Consider, for example, the Theory of Reasoned Action ,

proposed by Martin Fishbein and Izak Ajzen. Within this theory we might state as one

law, "the greater the attitude about the behavior, the greater the intention to engage in

the behavior." All the other paths implied by the diagram would be listed as laws within

the set of laws that define the theory of reasoned action.

2. The axiomatic format expresses relationships as a set of axioms. For example, within

the theory of reasoned ac tion, we might state as one axiom, "If attitude toward the

behavior, then intention toward the behavior." All the other paths implied by the

diagram would be listed as axioms within this format.

3. The diagram shown for the theory of reasoned action represen ts the causal

statement form. Each diagrammed path represents a theoretical proposition. For

example, we might infer from the diagram of the Theory of Reasoned Action that,

"the greater the attitude about the behavior, the greater the intention to engage in

the behavior."

Note Regarding the Format of Theory

The typical format used in sociology to express a theory is the set of causal statements,

often shown in a concise manner by the use of a diagram. In the 1980's, as part of an

effort to make sociology "more scientific," sociologists began to present their theories in

axiomatic format (see volumes of The American Sociological Review for examples of this

effort). Sociologists learned quickly that th e formatting of a theory provided few

advantages toward accumulating a scientific body of knowledge; what mattered was the

quality of the theory, not its formatting. Note, however, that some sociologists will argue

that "theory" should be expressed either as a set of laws or in axiomatic format (see:

Formal Theory in Sociology: Opportunity or Pitfall? , edited by Jerald Hage).

Truth of Statements, Validity of Reasoning

Peter Suber, Philosophy Department, Earlham College

True Premises, False Conclusion

0. Valid Impossible: no valid argument can have true premises and a false conclusion.

1. Invalid

Cats are mammals.

Dogs are mammals.

Therefore, dogs are cats.

True Premises, True Conclusion

2. Valid

Cats are mammals.

Tigers are cats.

Therefore, tigers are mammals.

3. Invalid

Cats are mammals.

Tigers are mammals.

Therefore, tigers are cats.

False Premises, False Conclusion

4. Valid

Dogs are cats.

Cats are birds.

Therefore, dogs are birds.

5. Invalid

Cats are birds.

Dogs are birds.

Therefore, dogs are cats.

False Premises, True Conclusion

6. Valid

Cats are birds.

Birds are mammals.

Therefore, cats are mammals.

7. Invalid

Cats are birds.

Tigers are birds.

Therefore, tigers are cats.

The distinction between truth and validity is the fundamental distinction of formal logic. You

cannot understand how logicians see things until this distinction is clear and familiar.

Simple statements

p "p is true" assertion

~p "p is false" negation

Compounds and connectives

p q "either p is true, or q is true, or both" disjunction

p · q "both p and q are true" conjunction

p q "if p is true, then q is true" implication

p q "p and q are either both true or both fal se" equivalence

Implication statements (p q) are sometimes called conditionals, and equivalence statements (p

q) are sometimes called biconditionals.

A truth table is a complete list of the possible truth values of a statement. We use "T" to mean

"true", and "F" to mean "false" (though it may be clearer and quicker to use "1" and "0"

respectively).

For example, p is either true or false . So its truth table has just 2 rows:

p

T

F

But the compound, p q, has 2 components, each of which can be true or false. So there are 4

possible combinations of truth values. The disjunction of p with q will be true as a compound

whenever p is true, or q is true, or both:

p q p q

T T T

T F T

F T T

F F F

If a compound has n distinct simple components, then it will have 2

n

rows in its truth table.

The truth table columns that define the basic connectives are as follows:

p q ~p ~q p q p · q p q p q

T T F F T T T T

T F F T T F F F

F T T F T F T F

F F T T F F T T

Most statements will have some combination of T's and F's in their truth table columns; they are

called contingencies. Some statements will have nothing but T's; they are call ed tautologies.

Others will have nothing but F's; they are called contradictions. Obviously these three types of

propositions exhaust the possibilities for statements that have truth table columns --which means

for all truth-functional statements.

DEDUCTIVE SYSTEMS

DEFINITIONS

AXIOMS AND ASSUMPTIONS

LEMMA AND LOGICAL RELATIONS DERIVED

THEORIES

COROLLARIES AND CONSEQUENCES

TIDItems

1Bread, Coke, Milk

2Beer, Bread

3Beer, Coke, Diaper, Milk

4Beer, Bread, Diaper, Milk

5Coke, Diaper, Milk

TID Items

1 Bread, Milk

2 Bread, Diaper, Beer, Eggs

3 Milk, Diaper, Beer, Coke

4 Bread, Milk, Diaper, Beer

5 Bread, Milk, Diaper, Coke

TID Items

1 Bread, Milk

2 Bread, Diaper, Beer, Eggs

3 Milk, Diaper, Beer, Coke

4 Bread, Milk, Diaper, Beer

5 Bread, Milk, Diaper, Coke

ItemCount

Bread4

Coke2

Milk4

Beer3

Diaper4

Eggs1

Itemset Count

{Bread,Milk} 3

{Bread,Beer} 2

{Bread,Diaper} 3

{Milk,Beer} 2

{Milk,Diaper} 3

{Beer,Diaper} 3

Itemset Count

{Bread,Milk,Diaper} 3

EXAMPLE: W e look at variables related to climate change and health concerns:

X1: average typhoons entering Philippine area of responsibility

X2: average annual temperature

X3: incidence of dengue cases per thousand population

X4: incidence of diabetes per thousand population

X5: incidence of tuberculosis per thousand populatio n

Over the last twenty years. The data are shown below:

typhoon avtemp dengue diabetes TB

26 33 15 13 12

25 34 14 12 13

15 29 6 11 5

13 30 14 14 12

8 34 12 10 11

11 32 14 12 13

25 26 17 14 12

9 35 8 11 7

18 31 14 11 5

15 32 10 13 8

Let us discover some association rules here. We use MINITAB to convert and simplify our data as

follows:

X1: Put a 1 if no. Of typhoons is more than 22 per year

X2: put a 1 if average temp. Exceeds 30

X3: Put a 1 if no. Of dengue cases exceeds 10

X4: put a 1 if no of diabetes cases exceeds 8

X5: put a 1 if no of TB cases exceeds 11

The converted data are shown below:

TYPH TEMP DENG DIAB TUBER ITEM

1 1 1 1 1 TYPH,TEMP,DENG,DIAB,TUBER

1 1 1 1 1 TYPH,TEMP, DENG, DIAB, TUBER

0 0 0 1 0 DIAB

0 0 1 1 1 DENG, DIAB, TUBER

0 1 1 1 0 TEMP, DENG,DIAB

0 1 1 1 1 TEMP, DENG, DIAB, TUBER

1 0 1 1 1 TYPH, DENG, DIAB, TUBER

0 1 0 1 0 TEMP, DIAB,

0 1 1 1 0 TEMP, DENG, DIAB

0 1 0 1 0 TEMP, DIAB

TID ITEM

1 TYPH,TEMP,DENG,DIAB,TUBER

2 TYPH,TEMP, DENG, DIAB, TUBER

3 DIAB

4 DENG, DIAB, TUBER

5 TEMP, DENG,DIAB

6 TEMP, DENG, DIAB, TUBER

7 TYPH, DENG, DIAB, TUBER

8 TEMP, DIAB,

9 TEMP, DENG, DIAB

10 TEMP, DIAB

SET OUR MINIMUM SUPPORT AT : M INSUP = 6 (ONE MORE THAN HALF THE NUMBER OF cases)

ITEM 1 SET SUPPORT

TYPHOON 3

TEMPERATURE 7

DENGUE 7

DIABETES 10

TB 5 (eliminate TYPHOON AND TB)

ITEM 2 SET SUPPORT

TEMP, DENG 6

TEMP, DIAB 7

DENG, DIAB 7 (ALL TWO ITEM SET ARE ACCEPTED)

ITEM 3 SET SUPPORT

TEMP,DENG, DIAB 5 (ELIMINATE THREE ITEM SET)

age gender salary exercise diet parents’history stress level

52 1 37664 3 5 1 21.4888

31 1 23400 1 2 2 12.1500

29 0 4000 3 4 1 6.4177

27 1 39379 4 4 2 22.2606

60 0 7000 1 2 2 6.1745

52 0 11500 1 5 2 9.6898

45 0 39715 1 4 1 21.2718

28 1 14344 3 1 1 9.0148

42 0 7229 2 3 2 6.5931

39 1 14609 2 2 1 8.3541

52 1 23678 1 2 1 12.6951

32 1 23758 3 2 0 13.3311

38 0 35271 4 4 1 21.1320

39 0 26000 3 2 0 14.9530

48 1 24805 2 2 0 13.6223

34 1 38579 3 3 2 20.5406

46 1 18000 5 5 0 18.1015

26 0 13864 1 2 1 7.7588

42 0 17119 4 4 1 12.5436

32 1 9516 2 5 1 7.9422

TYPE MGT NG0 INC.

1 5 0 4

1 4 0 4

1 5 0 5

1 4 0 3

1 3 0 3

1 4 0 5

1 5 1 5

1 4 1 4

1 3 0 4

1 3 1 3

0 2 0 2

0 2 0 2

0 3 1 5

0 4 1 5

0 2 0 2

0 1 0 2

0 5 1 1

0 2 0 1

0 2 0 2

0 1 0 2

The regression equation is

INC. = 1.57 + 1.16 TYPE + 0.272 MGT + 0.594 NG0

Predictor Coef SE Coef T P

Constant 1.5690 0.7091 2.21 0.042

TYPE 1.1648 0.5 792 2.01 0.046

MGT 0.2720 0.1 185 2.29 0.020

NG0 0.5939 0.6358 0.93 0.364

Unusual Observations

Obs TYPE INC. Fit SE Fit Residual St Resid

13 0.00 5.000 2.979 0.522 2.021 2.05R

17 0.00 1.000 3.523 0.734 -2.523 -3.01R

agency level LET

1 2 78

1 2 77

1 2 78

1 1 64

1 3 81

1 3 82

2 3 66

2 4 65

2 3 78

2 2 65

2 3 76

2 2 80

2 3 80

3 4 90

3 3 88

3 2 78

3 1 73

3 3 89

3 2 80

3 3 89

3 4 91

3 3 90

3 4 92

3 2 78

151312109871161434251

77.26

84.84

92.42

100.00

Similarity

Observations