Showing posts with label project. Show all posts
Showing posts with label project. Show all posts

Thursday, December 27, 2012

A statistical approach to predicting my getting sick

During my recent trip to Iceland, I thought to myself how grateful I am that I haven't been sick, and that it's surprising because "it's about time I got sick again."  Days later, I randomly got sick (sore throat, constant 102 temp, coughing); I should be fine in 3 more days, though.  In the meantime, just for fun during the winter break, I decided to see if I can define some formalism for this supposed intuition.

Disclaimer: I know this is some rough, make-shift calculations, but given the nature of having so little data, and that the real, useful data is all immeasurable and not at my disposal, this is what you get :-)

I get a sore throat about 5 times each year -- which always lends me voiceless for about a week -- despite the fact that I:
  • eat healthy
  • workout regularly
  • have good hygiene practices (shower at least once a day, brush teeth & floss twice a day, wash hands frequently, etc)
  • have no known allergies or health issues (have seen several docs at MIT regarding this)
I've been tempted to just go to NYC and lick every bus pole and escalator handrail I can, in attempt to go ahead and get it over with by subjecting my body to every known bacteria and virus out there.

Obviously, there's a myriad of variables that can cause one to get sick:
  • stress
  • germs being spread from air, physical contact, etc
  • weakened immune system (i.e., from over-training in the gym, stress, etc)
Clearly, these things are easily intertwined and impossible to accurately monitor and model.

However, in late 2007, I wanted to get to the bottom of it and was curious if there were any patterns in my getting sore throats.  So, for the past 5 years, I've quickly logged every time I get sick (a rating of the severity, how many days, and symptoms).

In the following picture, I plot every time I've gotten a sore throat, rounding to the closest week.



Looking at any patterns over time will only show possible correlations, not causes, I know.  Yet, the times I get sick are highly consistent from year to year:
  • I pretty much always get sick on Jan 1
  • I essentially never get sick from April 1 - July 31
  • my sore throat always lasts exactly 6 or 7 days.
  • I get sick roughly 5 times every year from August 1 - March 31.
  • 100% of the times that I fly overseas, I either get sick while there or soon as I come back -- but it's not always sore throats.  Here, I'm only concerned with sore throats.
I will refer to August 1 - March 31 as the sick period/months, which spans 32 weeks.  For all calculations, we are only concerned with getting sick during the current sick period, and that each one is independent of the previous years'.

Again, since it's impossible to model the aforementioned variables that are the actual causes of getting sick, I figured why not play around with the time-series data and treat it as indications as to when I'll get sick.  So, here we are making the large assumption that all of the underlying sick-causes variables are consistent and uniformly distributed during the sick months, and that I get sick on average 5 times per sick period.  I think this is reasonably fair as a high-level approximation.

If we model my sickness as a Poisson distribution, we see that the probability of getting sick N number of times during a sick period is:

N
probability of getting sick N times
1
0.033689734995427
2
0.084224337488568
3
0.140373895814281
4
0.175467369767851
5
0.175467369767851
6
0.146222808139876
7
0.104444862957054
8
0.065278039348159
9
0.036265577415644
10
0.018132788707822

I like that distribution; it seems to fit my sickness history pretty well... although places a little too much weight on getting sick more than 5 times.

For each one of these possible outcomes (and including all N up to however many weeks we have left... not just N >= 10), there's a chance of getting sick next week, which is simply N/(# of weeks remaining).... again, assuming uniform distribution of everything, including the times I get sick within the sick period.

Further, if we wish to calculate the overall probability of getting sick next week (i.e., summed over each possible N), we simply combine these two things, and weight and normalize by the Poisson probabilities listed in the above table.



As a real life example, last week:
m = 13 (meaning 13 weeks left until March 31 aka end of 'sick period')
n = 2 (meaning I've been sick twice so far)
lambda = 5 (meaning on average I get sick 5 times each year)

Per the above equation, the probability of getting sick this week (which I did) was roughly 42.10%.

Further, if I did not get sick this week, the probability of getting sick the following week would be 45.43%

Further, not getting sick in the 2 upcoming weeks, but getting sick on the following would be 49.14%

Thus, the probability of NOT getting sick during the current week and upcoming 2 weeks was only 16%, which agrees with irking feeling I got: "sweet; I'm not sick... it's been a while though..." 

In the meantime, if anyone is down for my NYC-immunity-deplete-boost idea, let me know.  I'll let you go first as the guinea pig.




Monday, February 20, 2012

Producing Ultra-Sharp Images

A few months ago, I bought a Zeiss 50mm f/2.0 Makro-Planar T* ZE Lens. This lens is hailed as one of the sharpest lenses in the world, and it's strictly a manual focus lens (no automatic focus).

Over the past year or so, I've taken 1,000+ pics of the jewelry that Sheley makes and sells. It's a learning experience, but clearly it's desirable to take pics that are incredibly clear and sharp.

So, with my new lens, and my side-hobby task of taking pics for Sheley, I was lead to the quest to produce what I'll call "ultra-sharp" images. I use this technique whenever I can; not just for jewelry photos.

My approach is to:
  • use a really sturdy tripod (my Benro A-1690)
  • use my manual focus Zeiss 50mm lens
  • I replaced my camera's internal focusing screen with the Eg-s screen, which assists my vision through the viewfinder by making the in-focus area appear very crisp and everything else really blurry
  • use the LCD screen at 10x magnification to view the subject, to help focus
  • use a remote shutter release (or put it on a 2-second timer) so that my pressing the shutter button doesn't shake the camera at all
Basically, it's an exciting process and requires much deliberation, patience, and carefulness for each shot. It's the closest I'll ever come to performing surgery.

Sample 1 -- The Penny


The viewfinder is so sensitive that moving the camera's focusing ring the slightest amount I possibly can is the difference between "In God We Trust" being in focus versus "2010" being in focus.

Sample 2 -- Boston Skyline 1/2 mile away


Notice how you can see each window with at least some detail. Pretty crazy considering the windows are half a mile away.

Sample 3 -- huge panorama of Boston from .8 miles away


Sample 4 -- Sheley's Jewelry

Unfortunately, images displayed on the web are rarely greater than 900 pixels in length, and are compressed greatly; thus, the increase in sharpness may often not even be discernible for images posted to the web. Oh well, it's been really fun for me to discover how to push my camera and lens to their limits.

In the meantime, I've become a huge fan of Zeiss, and I can easily understand why they're considered a world leader and pioneer for producing world-class optical equipment, from medical lenses, microscopes, camera lenses, etc.

Saturday, October 16, 2010

my car gets more character, and by character i mean ghetto

(note: if you are viewing this in facebook, the 4 images may not appear. click refresh until they do, or visit the original source: http://omggetthis.blogspot.com)

i've never placed much stock in any car i've had. my old '88 honda accord was quite amazingly ghetto/enjoyable though. it had a leaking sunroof which would sometimes pour in tons of rain after driving it for the 1st time after a big rain. sometimes it wouldn't. it was always a surprise. it had other quirks too, of which i'll spare you.

my current car is a '99 nissan altima, and we've been building a relationship for 3 years now. here it is, street-parked next to my boston apartment (this is 1 of 2 times in months that i havent had to parallel park):

most notable are our following adventures: (1) 50 hours of fall foliage driving; (2) driving from boston to atlanta and back, while almost dying in a multi-hour snow storm w/ no windshield wipers -- forcing me to drive on highway w/ head out the window; (3) car broken into twice...

and now, the latest addition:
the trunk release broke many months ago. my trunk key has never worked. so, i've had no access to the trunk for months. i fiddled around w/ it and learned how the trunk works.

if you slide this metal piece towards the right, the trunk will open.

i tied string around the sharp metal, while adding padding to prevent the string from unraveling.

i ran the string along the base of the trunk, bended it through a loop, then ran it through the back seats towards the front of the car. now, if you pull the string towards the front of the car, the trunk will open.

here i am in the driver's seat, turned towards the back.


it's nothing complicated; i just think it's kind of a funny system.

Sunday, June 27, 2010

student loans: my trick on how to pay them off faster

regarding my loans, i saved roughly $250 within a 12-month span. you can too.

short version:
i wrote a little java program that uses simulated annealing to find the near-optimal solution of how to distribute your monthly loan payment amongst your various loans every month, while minimizing the amount of accrued interest.


the above .zip file contains 3
files:
for example, run it from a command prompt:
>> java InterestOptimizerTest input.txt


long version:

if you have many loans which total a lot of money, you may be waisting 100s of dollars merely by incorrectly paying the wrong amounts towards each loan. the point of this writing is to help you save money.

obviously, one can save/waste a lot of money depending on when and how much they pay off their loans. this isn't about that. say you have 6 loans and $1,000 you wish to pay towards them. this writing is about how to apply your $1,000 towards the 6 individual loans in the best possible way so that you aren't wasting money. interest is a killer. below, i mention my approaches towards minimizing the interest, and i provide my actual student loans as an example:

upon graduating from ucla last year, my student loans totaled roughly $78,000. it's been 13 months since i've graduated, and i've gotten them down to $27,000. here are my 6 outstanding loans:

loan name, balance, APR, minimum monthly payment
-------------------------
sallie mae 1, $13472.75, 2.58%, $201.03
sallie mae 2, $1544.54, 7.25%, $214.74
sallie mae 3, $5032.22, 6.8%, $221.29
doe 1, $521.06, 8.5%, $52.61
doe 2, $2378.14, 6.8%, $59.50
ecsi, $4495.36, 5.0%, $60.46

for the rest of this article, let's be consistent and say that i pay $2,000 each month.


idea #1 - naive approach based on distribution of annual interest:
this is what i had been doing the whole time. when i had to start paying back my school loans, i didn't give it much thought as to how i would determine how much to spend towards each. i just quickly thought, 'i'll just pay some amount that is proportional to how damaging the interest is.' i now regret not thinking about it more.

the amount i was paying towards each loan was merely determined by its contribution towards my total annual interest across all loans:

loan i's payment := $2,000 * (loan i's yearly interest / total yearly interest across all loans)



idea #2 - brainstorming + linear/non-linear programming:
after filing my taxes, i noticed i paid a lot towards loan interest this past year. so, while brushing my teeth in the morning, i thought to myself that my current method (idea #1) is probably much worse than what i imagined. how can i find the optimal solution? first thing that popped into my mind was to search the solution space by using a genetic algorithm, but by the time i finished brushing my teeth, i had convinced myself that the sheer thought of such was complete overkill, laughable, and non-optimal, and clearly not the best approach.

throughout the day, i thought about other approaches:

linear/non-linear programming:
the idea of using linear programming came to mind. the format would be:

minimize C^T * x || sum(x) = b, where
  • C is a vector of known coefficients that represent the interest for each loan
  • x is a vector we are solving for, which represents how much to spend on each loan
  • b represents how much we are willing to spend in a given month (i.e., $2,000)
in other words, we want an equation for our total interest across a given time period (1 month), and we want to minimize it. but unfortunately, we can't really form a linear equation for it. for example, the interest for one loan would be:

loan i's interest = (loan_i's_original_balance - how_much_to_pay_towards_loan_i) * (1 + APR/365.25)^30

we are minimizing the sum of this value across all loans.

that exponent makes it non-linear. moreover, i'm not sure how to solve it using non-linear programming either, as we no longer have a clean coefficient. we don't have an Ax = b; we have f(x) = b

maybe there's some way to use the simplex algorithm for this, but i wasn't interested in finding out. seems messy.



idea #3 - simulated annealing:

well, when all else fails, use simulated annealing, right? haha. simulated annealing is a pretty simple, common AI approach that seeks to find an optimal solution within a large search space.

i implemented it in a little java program.

basically, it:
  • initially assigns random payment amounts towards each loan, while meeting the constraints of the minimum and maximum payment for each loan.
  • it then calculates the total monthly interest across all loans.
  • then, for however many epochs i care to run my algorithm, it will:

    • randomly decide to increase or decrease a to-be-selected loan,
    • then the loan it chooses to adjust is probabilistically chosen, based on each loan's contribution towards the total amount of interest across all loans.
    • the amount that it perturbs the loans is randomly chosen, too.
    • if the new payment amounts yield a lower interest, let's save these payments as the current state and continue
    • if the new payment amounts yield a higher interest, let's probabilistically save these payments as the current state and continue (this is to help avoid local minimums)
    • saves the best payment distribution that we encounter
the only tricky part was ensuring that the perturbing of the payment amounts still satisfied all of the constraints across every loan. a lot of error-checking tedium.

the program will display how much to pay towards each loan for every month, and it'll produce a summary that tells you how much you paid total, and how much went towards interest.

you can also experiment with choosing different monthly payment amounts and see how much it affects your lifetime interest amount. for example, my program showed me that:
  • paying $2k/mo will take 13 months; i'll accrue ~$500 of interest.
OR
  • paying only $1k/mo will take 26 months; i'll accrue ~$1,000 of interest
so, if i take a bullet of paying a measly extra $500 in lifetime interest, i'll gain the
freedom of having $1k extra of non-loan money each month for the next 26 months! it's definitely worth it.



solution:

the above .zip file contains 3
files:
for example, run it from a command prompt:
>> java InterestOptimizerTest input.txt




future work:
my current solution attempts to minimize the accrued interest, but it only does so over each upcoming month. basically, it works in a month-to-month basis until all loans are gone, which is not guaranteed to be the optimal, global solution, since it's not looking at the bigger picture across the lifetime of all the loans. the optimal solution would need to apply random amounts to all loans across every month of its lifetime, then make shifts both (1) between the various loans and (2) from month-to-month. this would give it a global picture to search within.

i suspect the improvement would be relatively negligible. in my case, i'm guessing that it would save me no more than $20. so, i don't plan to add this improvement. it's not worth it, but yea, i recognize it as an area of improvement.

ideas?
if anyone has any ideas for a faster or more optimal solution, please let me know. i'm curious to hear suggestions.