 |
|  |
 |
|  |
 |
|  |
 |
|  |
 |
|
DrSpike
|
 |
Enthusiastic member of Apolyton
Sep 2001 time: 05:25
|
|
I build and estimate cutting edge econometric models for a living. I have a ph.d in econometrics (statistical economics). However, screw that, your comment about my expertise is not what bothers me. What bothers me is your reaction to my suggestions........it really doesn't seem in keeping with the picture I have built up of you.
Still, because I respect you, one last try:
The problem with the control you propose here is that it itself is noisy from your proposed sample.........a far better control would just be to use the formula we know holds (and the formula for bog standard comabt is the same for MP, we can be virtually certain of this - it is just the bonus modifiers that may be different against a human rather than an AI).
As to the arbitrary 18, the point is not that your educated guess is a bad one, but that you need a concrete decision rule. Your statement:
"I'll bet, on the first surprise attack, that the attacking unit will win 18 or more of the 25 encounters, and 22 or 23 won't surprise me. I'll be real interested to see if the results of attack 3 match attack 1 or if they're consistent with 2 and 4"
Now 'consistent' is the key bit........how much of a divergence are you going to allow before it affects your conclusions? Do you know the probability of the warrior winning more than 18 even if a bonus does not exist. No, you don't. Do you know the probability of observing 17 when a bonus does exist? No, you don't. Both of these are ways your decision rule fails.......now potential failure in hypothesis testing is endemic, but you have to know the associated probabilities or results are just glorified junk.
It's funny, I always stay out of these stats debates, because it is far easier to let the person do it by brute force, with a badly constructed experiment but lots of repititions than try and explain a more cogent testing structure. But you said you were going to talk to stats people about confidence, so I chirped up, and thought you would be receptive.
Still, happy testing........hope it goes well.
|
|
|  |
 |
|
rah
|
|
Apolyton Prince of Moderators, Master of Reason
|
 |
Lord of the Ferrets
Jan 1970 time: 23:25
|
|
I guess I overreacted to your belief that human vs humanl combat need not be tested and that previous ai vs human results would be valid despite never having been tested. Hence my desire to have a control. (but I do agree that they will probably be the same, but paranoia when testing is never a bad thing)
And your comment that 18 was arbitrary. It was not.
An educated guess is quite different than arbitrary.
And to your nitpicking on 17 or 19. WE DON"T KNOW WHAT IT"S GOING TO BE, that's why were testing.
The initial hyposthesis is that there is a difference.
And I'll do the statistical run, AFTER the test. Having data has been known to make it easier. 
Since the testing will consume some time, I wanted a statistician to give me a ballpark of number of observations needed so I wouldn't not get enough or waste my time doing hundrends.
If I overreacted, I apologize, but I recommend that in the future, if you don't want people to get defensive, don't start by attacking there initial position distorting what they said. And your use of abritrary and misinturperting my hypothesis (there is a difference, the amount of difference was just a guess) was exactly that.
RAH
|
|
|  |
 |
|
DrSpike
|
 |
Enthusiastic member of Apolyton
Sep 2001 time: 05:25
|
|
quote: Originally posted by rah
And your comment that 18 was arbitrary. It was not.
An educated guess is quite different than arbitrary.
RAH |
Again, sure you used knowledge to select 18, but you don't know any of the properties of that decision rule or you would have posted them to shut me up by now. Maybe arbitrary is a strong word, but it's not far off. 
I doubt anyone really cares, but this is how I would do it. Far from being clueless in this area as you suggested I do this (well this is analogous to the most basic econometrics I would do) every day, and lecture other people at undergraduate and postgraduate level on how to do it.
I would use 50 repitions (or 25*2 as you suggest is the same, as long as making and breaking peace gives you the same state as the initial one before the civs have met). I would derive the proportion of expected wins for the attacking warrior without a bonus. Your null hypothesis is then.
H0: p hat = derived number
with
H1: p hat is not derived number.
Then you just derive the variance of phat in your sample of 50......it is just {p(1-p)}/n, where n is the number of repititions Under the null the distribution of phat is normal (from various statistical theorems) with a mean of 'derived number' and a variance as above. You then choose a level of significance, which also fixes the probability of you rejecting the null when it is true.
You then compare the standard normalised value for the proportion of wins you observed in your testing with the associated critical value from the normal distribution, which tests the null at the 5% level, the level you said you wanted.
Easy. Concrete. And as the icing you can derive the power of the test (related to the probability of accepting the null when it is false) under the assumption of a bonus of 50%. YOU CAN EVEN DERIVE THE MOST LIKELY SPECIFICATION FOR THE BONUS GIVEN THE DATA.
Please feel free to show this post to your stats guys if you doubt my expertise as you said above. Any competent statistician would carry out the test the same way. 
Last edited by DrSpike on 31-01-2003 at 21:31
|
|
|  |
 |
|
rah
|
|
Apolyton Prince of Moderators, Master of Reason
|
 |
Lord of the Ferrets
Jan 1970 time: 23:25
|
|
quote: Originally posted by DrSpike
I would use 50 repitions (or 25*2 as you suggest is the same, as long as making and breaking peace gives you the same state as the initial one before the civs have met). I would derive the proportion of expected wins for the attacking warrior without a bonus. Your null hypothesis is then.
|
I'm doing 25*2 to test if the different states make any difference. (can you get a second suprise bonus) But I'm glad you mentioned that because I was going to just line up 4 warriors against 4 warriors, now after reading this, yes the first combat must be run right after the initial contact notice, So I'll start with the four warriors 1 square away as the intial base point (i probably would have remembered that when I set it up but I might have wasted time ).
As to your statitiscal analysis. Yeah I could have copied crap from my SAS book, but I have statisticians working for me and their expertise there is greater than mine. I have never once claimed expertise there, just understanding. My expertise is the test methodology and data collection/manipulation. Analysis comes after the data is collected. And if you want, I'll give you the raw data and you can save my staff some effort
But, if the result is > 18, I'll claim there is a bonus prior to the stat run, but I'll wait for the analysis to tell me how much and how confident it is. My initial intention was too disprove all of those that claim that the sneak attack bonus is a myth. How much has always been secondary.
And the difference in win percentage is all MP players are really interested in so they can develop some rule of thumb guidelines for on the fly attacks. When the clock is ticking all you really have time for is hmmm 4attack vs 2def, good, or 4att vs 4 def, I'd better have more units attacking than he has defending. or 4att vs 6def, I'd better have more than twice the units attacking.
RAH
|
|
|  |
 |
|  |
 |
|  |
 |
|  |
 |
|  |
 |
|  |
 |
|  |
 |
|  |
 |
|
rah
|
|
Apolyton Prince of Moderators, Master of Reason
|
 |
Lord of the Ferrets
Jan 1970 time: 23:25
|
|
warriors, for two reasons.
One, it takes faster to set up on a random start, I'll just do it on a small world and it won't take war and i more than 10 minutes to line em up, and i won't have to use a scenario that could bias the results.
Two, my initial intent was just to prove there was a bonus. so any equal units would do. IF there is one, we'll work on the amount.
With Spike, we're fine, we were just feeling out each others backgrounds, and I took early offense to a few of his choice of words, and comments about not needing a control group. I overreacted and fueled the flames. My initial thought was a simple test, but it's kinda gotten bigger. I believe we're fine and will share on the discovery. I'm glad there's a statistician available since I really don't like using my staff for my personal concepts(even though I know they'd be happy to do it) This way we can keep the entire thing in house. 
RAH
And everyone showed up early to play so I didn't get a chance to test it, War and I will do it tomorrow before we play.
|
|
|  |
 |
|  |
 |
|  |
All times are GMT. The time now is 05:25. Apolyton Time is 00:25. |
top of page
|
| archivepost |
|
Forum Rules:
You may not post new threads
You may not post replies
You may not post attachments
You may not edit your posts
|
HTML code is ON
vB code is ON
Smilies are ON
[IMG] code is ON
|
|
|
|
|
|