Monday, July 9, 2018

Understanding the Statistics of Kalman Filters I

Kalman Filters uses the concept of correlation and regression. I use the following example.
Let use predict the position of car at time k on the basis of position of car at time  k-1 and velocity of the car at time k. 
So, we apply the concepts of simple linear regression. Here A is the change in the position of car at time k with a unit change in the position of car at time k-1. B is the change in the position of car at time k with a unit change in the velocity of car at time k. Further, 
As these two linear models mentioned above approximate a real life scenario some amount of error is always involved. This error is represented by wk and vk. These are independently and identically normally distributed with a mean 0.These errors are called measurements errors or white noise. This discussion will be continued in the next entries of this blog.

Friday, February 9, 2018

Why do we need to "Test the Hypothesis"?

Testing ESP

The  following example (metaphorically and physically) illustrates the power of the "concept  of hypothesis testing" in introducing objectivity  in any type of research. This could range from social science to medicine. This concept is explained by an example on Extra Sensory Perception (ESP).  A  person  with  ESP is normally is looked at with suspicion as ESP cannot be validated with exact science with measurements and readings. But with "Testing of Hypothesis" ESP or Sixth Sense can be tested or validated objectively. 
Suppose there are 50 red and black cards and a person guesses the color of 32 cards out of 50 cards correctly , in an experiment. In this experiment the person with ESP is supposed to guess the color of cards correctly. The cards are displayed in one room and the person with ESP is in next room. He/she has no prior knowledge of the color arrangement of the cards.
Under normal situations if a person guesses the color of 50 cards out of 50 cards correctly, he is supposed to have some unsual power of foresight called ESP. Also if a person guesses 0 out of 50 cards correctly, he is supposed to have no foresight or no ESP.  Now we try to answer the following  important question statistically.
OUT OF 50 CARDS HOW MANY CARDS SHOULD BE GUESSED CORRECTLY IN ORDER TO BE 99% SURE THAT THE PERSON HAS ESP?

Statistical Illustration of the Example

Statistical Solution to ESP Example



Friday, October 13, 2017

Markov Chains and Markov Process - Part III


Markov Chain and Markov Process III unveiled in my lecture video available in this link <https://youtu.be/WQUaLNK7QAc>. Please visit this link.

Thursday, October 5, 2017

Markov Chain and Markov Process - Part II

Taking a sick leave

The discussion on markov chain and markov process as a special case of conditional probability is continued today with an example of sickness and good health. The state of good health or bad health tomorrow for an individual can be explained by Markov Chains and Markov Process in the following manner.
P(bad health tomorrow )
=P(bad health today)P(bad health tomorrow|bad health today)+P(good health today)P(bad health tomorrow|good health today)
If a person is ill today  then the probability that he/she will feel ill tomorrow is higher than the probability that the person with good health feeling ill tomorrow. 
This can be explained by following notation.
P(bad health tomorrow|bad health today)>P(bad health tomorrow|good health today)
What is the probability that a person is ill on the third day (of the week)?
P(Bad health on third day)=P(bad health on second day)P(bad health on third day|bad health on second day)+P(good health on second day)P(bad health on third day| good health on second day) 
This  can be interlinked (like a chain) to the following expression
P(Bad health on second day)=P(bad health on first day)P(bad health on second day|bad health on first day)+P(good health on first day)P(bad health on second day| good health on first day) 
This inter-linkage (chain) will be discussed in detail in the next blog. 

Friday, September 15, 2017

Markov Chains and Markov Processes

Here the concept of conditional probability discussed in previous days  blogs (July 14 - July 30, 2017) is extended to Markov chain and Markov Processes. 
Let,
P(Rainy Day) = 0.3
P(Dry Day) = 0.7
We know that if it rains today then there are high chances that it might rain tomorrow that is the probability that it will rain tomorrow is high.  So the conditional probability that it will rain tomorrow given that it rains today is higher than the unconditional probability that it will rain tomorrow.
Notation wise
P(Rain Tomorrow|Rain Today)>P(Rain Tomorrow)
Conditional Probability>Unconditional Probability
P(Rains tomorrow)=P(Dry Today)P(Rains Tomorrow|Dry Today)+P(Rains Today)P(Rains Tomorrow|Rain Today)
Markov Processes are systems where outcome of the current state is highly dependent on the outcome of immediately preceding state.  Examples can be weather systems and state of well being of an individual. Here outcome of state today is heavily dependent on the  outcome of immediately preceding state. So probability that it rained one month ago has less influence on the probability that it rained today than the probability that it rained yesterday. State of well being of an individual can also be explained in the following manner. So these examples explain special case of conditional probability namely markov chains. This will be discussed in detail in the next blog.

Wednesday, August 23, 2017

Expectation and Conditional Expectation

The discussion on probability and conditional probability (previous day's blog) is continued to expectation and conditional expectation.
Let Z be a random variable denoting money spent by the state on the cardiac care of a citizen below 40 years.
E(Z) is the average money spent by the state per citizen below 40 years on its cardiac care.
E(Z)=P(Y1)E(Z|Y1)+P(Y2)E(Z|Y2)
=(Probability of a Heart attack before 40)(Average money spent by the state on caradiac care of a citizen with an incidence of heart attack before 40 years)+(Probability of no heart attack before 40 years)(Average money spent by the state on cardiac care of a citizen with no incidence of heart attack before 40)
so,
Expectation=Probability(condition1)* Conditional expectation+Probability(condition2)* Conditional expectation

Sunday, July 30, 2017

Probability and Conditional Probability


The discussion is being continued. 
So,
Y1 stands for  incidence of a heart attack before 40.
Y2 stands for no incidence of a heart attack before 40.
X stands for the event of Blood pressure and Blood sugar beyond normal limits and cholesterol within normal limits.

Expenses on Cardiac Care


Probability and Conditional Probability


Wednesday, July 26, 2017

Probability and Conditional Probability

This is in continuation of the discussion on Probability and Conditional Probability of the Previous Blog.
Now we are using a general notation. Then the equation get's reduced to the following form.

Probability and Conditional Probability

Tuesday, July 18, 2017

Visualization of Probability & Conditional Probability


 
The tree diagram given in the image below helps visualize the main difference between conditional and unconditional probability. Here we continue with the example of the previous blog where conditions leading to a heart attack before 40 years are minutely analyzed.  

Visualization of conditional probability

Probability and Conditional Probability


Friday, July 14, 2017

Understanding conditional probability through Venn Diagrams

Conditional probability is normally computed under some additional conditions and "unconditional probability" is the probability where no additional conditions are provided. This is illustrated with example mentioned in the previous blog. Here the probability of having a heart attack before 40 years of age is minutely analyzed by a Venn Diagram. The venn diagram  below shows a rectangle representing all people ( in an area) below 40 years. The circles H, S, C and N  are those who have suffered a heart attack. The other region represents those without  a heart attack. The shaded region in the diagram below gives the  probability that a patient with  high level blood pressure and sugar but the level of cholesterol is within limits has a heart attack before 40 years of age. This is the ratio of number of elements in the shaded region divided by the total number of elements in H, S, C and N.

Friday, July 7, 2017

Understanding Venn Diagrams


Incidence of Heart Attack Below 40 years of Age

We are interested in finding the probability of a heart attack below 40 years of age. Venn diagrams help visualize the scenario of people having heart attack before age of 40 including those with high blood pressure denoted by H, high sugar denoted by S and high cholesterol denoted by C or a combination of these. N denotes those individuals who have stroke before 40 and have blood pressure, sugar and cholesterol within normal limits. The image below shows the venn diagram of those cases where the blood pressure and sugar are beyond normal limits but cholesterol is within the limits.  This way all the possible causes resulting to a heart attack before 40 years can be minutely analyzed and it's probability can be calculated. This analysis is based on collected data and can aid in correct diagnosis.



This venn diagram shows probability of a stroke with high levels of blood pressure and blood sugar


Tuesday, June 27, 2017

Exploring aic (Akaike Information Criterion)

Goodness of Fit of Statistical Models 

aic plays a key role in the interpretation of efficiency of probability models with respect to a model of reference. Here the ratio of two likelihood functions is taken and it is equivalent to the difference between the two loglikelihood functions. So the greater the difference between loglikelihood functions the greater is the difference in efficiency  between the models.
aic=-log(L(P1)/L(P2))+k
Here P1 is estimator of the parameter of model 1 and L(P1) is the likelihood function obtained from probability model 1. L(P2) is the likelihood function obtained from probabilitymodel 2, with P2 as the estimator of the parameter of this model. k is the degree of freedom.
aic = -[log(L(P1))-log(L(P2))]+k
This expression shows that the greater the difference [log(L(P1))-log(L(P2))] the higher the value of aic. So high value of difference indicates greater dissimilarity between model 1 and model 2. For relatively close models aic should be lower than the aic of relatively distant models. Smaller value of aic implies that the fit is good. The model with higher value of aic ( with respect to a model of reference)should be rejected. 

Saturday, June 24, 2017

Fertility Rates as Development Indicators

Vital Statistics as Development Indicators 

The discussion on fertility rates as development indicators discussed in the previous BLOG is continued here. In the image below the fertility rates of Nepal and Germany are compared. These fertility rates are namely  Age specific fertility rate (ASFR) and Total fertility rate (TFR). This comparison shows that the development levels of two countries can be compared by comparing their fertility rates. In the previous blog a comparison between TFR was done to explain this concept . Here in the table (image)below ASFR and TFR are compared. ASFR is defined as the number of birth per 1000 women in that age group. For example, the ASFR for the age group 20-24 for Nepal in the year 2000-2005 and 1995-2000 is 231.2 and 257.4, implying that out of 1000 women in the age group 20-24, 231.2 gave birth in 2000-2005 and 257.4 gave birth in 1995 - 2000. This number has declined sharply from 1995 to 2005. Whereas in Germany this value is 53.4 and 58.5 respectively. In Germany the decline in not so drastic as compared to Nepal. ASFR attains its peak in Nepal in the age group 20-24. This implies that in Nepal most of the women bear children in the age group 20-24.The ASFR  data of Nepalbased on census 2011also exhibits this pattern. But in Germany ASFR is highest in the age group 25-29. As the country becomes more developed not only does the TFR decrease (Germany 1.3 , Nepal 3.7 in 2000-2005) but the age at which ASFR attains its peak  increases. The reason behind this is that as the development activity increases more and more women become economically active through various job opportunities coupled with such developments, this results in decrease in TFR. Many women shift their age when they bear their first child to later years (25-29 in Germany) resulting in highest ASFR in 25-29.

Comparison between Germany and India


Monday, June 19, 2017

Fertility Rates as Development Indicators

Development Indicators


Commonly used fertility measures are Total Fertility Rate (TFR), Age Specific Fertility Rate (ASFR), Crude Birth Rate (CBR), Net Reproduction Rate (NRR) and Gross Reproduction Rate (GRR). The rate of growth of population is viewed from different perspectives through these measures. Total fertility rate (TFR) is defined as the number of children that a woman bears in her entire fertility span. According to 2011 census TFR of Nepal is 2.52 with 1.52 for urban areas and 3.08 in rural areas. This means that a woman bears 2.52 children in her entire life span. The women in urban areas bear 1.52 children whereas the women in rural areas bear 3.08 children. TFR of Nepal in 1971 was 6.32. This shows that Nepal has made a big progress in the reduction of TFR from 6.32 in 1971 to 2.52 in 2011. A low TFR is related to high average life expectancy and low Infant Mortality Rate (IMR). In 2011 Census TFR of 2.52 is coupled with 66.6 years of average life expectancy and 40.5 IMR, implying that in 2011 a woman bears 2.5 children in her life time, a child born has life of on average 66.6 years and there are 40.5 infant deaths per 1000 live births. In the absence of use of contraceptives and birth control measures the TFR of a country is around 6. This is reflected by TFR of Nepal in 1971 of 6.32. High TFR is coupled with high IMR and low life expectancy at birth. This fact is validated by the rural and urban differential in these measures. According to 2011 census of Nepal, TFR urban is 1.54 and TFR rural is 3.08, IMR urban is 24.06 and IMR rural is 42.9 and finally the life expectancy at birth of Urban areas of Nepal is 70.5 and 66.6 for rural areas. High TFR indicates poor health facilities and health conditions in governmental hospitals in rural areas in contrast to urban areas. This is also coupled with a lower value of average life expectancy in rural areas and a very high IMR. High IMR also implies large deaths of infants due to poor nutrition of mother and poor health facilities.  Many countries in Africa like Sierra Leone, Angola have high incidence of diseases like Malaria and HIV AIDS have low average life expectancy (50.1and 52.4 years respectively in 2016) and high TFR (4.76 and 5.31 respectively in 2016). Many countries with low TFR (lower than 2), have a declining population as two people mother and father are replaced by less than 2people.  Countries with low TFR have high average life expectancy and low IMR. For example in 2016, TFR of Italy is 1.43, this is very low and it is coupled by a very low value of IMR of 3.3 deaths per 1000 live births and average life expectancy of 82.7 years. This is due to good health facilities provided by the government of such countries. A low value of IMR also indicates good nutritive diet received by the mother. So TFR can be related to the development status of a country, where a country with low TFR has high socioeconomic and developmental status in contrast to countries with high TFR. This discussion will be continued in the coming BLOGS.  

Thursday, June 8, 2017

Maximum Likelihood Estimation illustrated with an example

Predicting the risk of Zika Virus infection

Here the concept of MLE explained in the blog on June 6 is explained with an example. Denoting by P the proportion  of babies infected by the Zika virus,  among the population of Zika virus infected pregnant mothers in an  area like say Brazil, this P is a population parameter. P is unknown, but we want to know its true value. But it is impossible to know the entire population and the population parameter due to time and monitory constraints. 10 samples of size 50 each are selected for the estimation of P. That is 10 samples of 50 Zika infected pregnant women are selected. So, n = 50 and m = 10. We are interested in finding the number of babies infected with zika virus among 50 zika infected pregnant women selected in 10 batches. So Xi is the random variable denoting number of zika infected babies in a sample of size 50. So, Xi = 0, 1, 2, ....50. So the MLE is the number of zika infected pregnant women bearing zika infected babies divided by 500. This gives the fraction of zika infected pregnant women bearing zika infected babies among 500 zika infected pregnant women and this is the MLE of P.

Tuesday, June 6, 2017

Exploring MLE and Maximum Likelihood Estimation

Maximizing the Likelihood

Maximum likelihood estimation chooses a sample statistic that maximizes the likelihood (probability function) of occurrence of sample for a particular parameter. Maximum likelihood estimation is based on the concept of Maxima and Minima. Through maximum likelihood estimator (MLE) an estimator is chosen; this estimator maximizes the likelihood of occurrence of the sample. The probability function of the occurrence of the sample is maximized by taking the derivative of the  log of the likelihood function (with respect to the parameter) and equating it to zero. This gives an estimator ( based on the sample) of the parameter. The second derivative of this loglikelihood function will be less than zero for this MLE. We also know that the probability function of this sample is based on a population parameter. Population is unknown and so is the population parameter. But the sample is known and we try to find the estimator based on the sample. This estimator maximizes the probability (likelihood) of this sample. For example 
Xi~B(n, P) that is X follows Binomial with parameter n and P. Here i = 1, 2, ....m
Then the MLE of P is  Sum of Xi (over all m)/nm and is given in the image below. This will be illustrated with an example in the next blog.

Tuesday, May 30, 2017

Rate or Ratio?

Mortality Rates

Vital Statistics analyzes vital events occurring in the life of an individual. Mortality, fertility and migration are some examples of such vital events. Vital statistics comprises of measures such sex ratio, crude birth rate, crude death rate etc. But terms Rate and Ratio are some what ambiguous. The definition of ratio here doesn't seem to confirm with the traditional mathematical definition which is as follows. A ratio is written as " a to b" or "a:b" and is expressed as quotient of two numbers.  Rate is mathematically defined as quantity measured with respect to another quantity. So rates are always defined in terms of  two units like miles per hour for speed or number per milliliter of water for bacterial growth. 
This ambiguity is illustrated with following example. Sex ratio is defined as number of males per 100 females.  So if sex ratio is 106 means by definition that there are 106 males per 100 females. So sex ratio 106 doesn't confirm with the traditional mathematical definition of ratio. Also if sex ratio is given as 1.06 then one should conclude that  there are 1.06 males per 1 female which is equivalent to the previous statement of 106:100. Now maternal mortality ratio (MMR) is defined as annual number of maternal deaths due to pregnancy related causes per 100,000 live births. But maternal mortality rate (MMR)is also number of maternal deaths per 100, 000 live births. Similarly Crude death rate (CDR) is number of deaths per 1000 people in the population and Infant mortality rates (IMR) number of deaths of infants per 1000 live births.

Intuitive approach is the best possible option for finding a way through this ambiguity. We should not be carried away by these numbers and loose our way in this ambiguity of definition. We should rather link these mathematical figures with country specific scenario. This way we will not loose track of the main objective, which is  getting a clear picture of development status of that country.
.........and one more tip in case of frequent events  like birth and death it is per 1000 and in case of less frequent  events like cause specific death rate it is expressed per 100,000.


Friday, May 26, 2017

Relation between Binomial, Poisson, Geometric and Negative binomial distributions

A small grocery store in a very busy market square

The interrelationship between these distributions is illustrated with following example.
Suppose that there is a small grocery store located in a very busy market square with several big stores and shopping malls. Hundreds of people are walking in this market square in a Saturday morning. Then the number of people entering this store is a random variable following Poisson distribution. Number of people making purchases upon their entry into this shop is a random variable  following Binomial distribution. Number of people making purchases before a person leaves the shop without buying anything is a random variable following geometric distribution. And lastly the number of people entering the shop before third person makes purchases more  Rs. 5, 000/- is also a random variable and follows negative binomial distribution. 
In this example, different random variables governed by probability law of different probability mass functions explain different aspects of a purchase in a small shop.

Sunday, May 21, 2017

Statistical Analysis and Modelling of Vital Events

Mortality and Fertility Models for Countries with Limited Data

My book titled " Mortality and Fertility Models for Countries with Limited Data" is available on Amazon.  

It elucidates in a simple manner various methods of data generation, correction, prediction,  analysis  and interpretation. Here  I discuss the statistical analysis of vital events  namely Mortality , Fertility and Migration for countries with limited data like Nepal and India. 
Various techniques of data correction, data analysis and data prediction can be learnt from this book. There are several examples of vital events from Nepal, India and Germany in this book.

Learn various statistical techniques of data correction, analysis and interpretation


Friday, May 5, 2017

Insights into the concept of hypothesis testing (last Part)

When alpha is 0.025(z = 1.96) beta is 0.1866 and when alpha reduces to 0.005 (z = 2.55) beta increases to 0.3538.

When alpha is 0.025

Alpha is 0.025 and Beta is 0.1861


Alpha is 0.005 and Beta is 0.3583


The mathematics mentioned above is explained in this image