20 September 2009

Simulation Models

In scientific computing, we often are looking to model physical processes through computer simulation. Though generally computational scientists (physical scientists that is) are performing this work, often this simulation and modelling is not done in a very systematic or scientific way. This has been an issue for computational science since the inception of the computer.

This challenge was recognized many decades ago and was the driving force in John Pople's computational quantum chemistry research. Though Pople's chosen field was quantum chemistry, we would do well to learn from his approach to modelling for our chosen fields of research and study.


Here in general terms and for an outline of development and use of a model (from Pople.)

A theoretical model for any complex process is an approximate but well-defined mathematical procedure of simulation.

......

Five stages may be distinguished in the development and use of such a model:

1) Target
2) Formulation
3) Implementation
4) Verification
5) Prediction



Now an explanations of these stages from Pople's quantum chemical models:

1) Target
A target accuracy must be selected. A model is not likely to be of much value unless it is able to provide clear distinction between possible different modes of molecular behavior. As the model becomes quantitative, the target should be that data is reproduced and predicted within experimental accuracy. For energies, such as heat of formation or ionization potentials, a global accuracy of 1 kcal/mole would be appropriate.

2) Formulation
The approximate mathematical procedure must be precisely formulated. This should be general and continuous as far as possible. Thus, particular procedures for particular molecules or particular symmetries should be avoided. If this can be done, the procedure becomes a full theoretical model chemistry, which can be explored in detail as far a available resources permit.

3) Implementation
The formulated method has to be implemented in a form which permits its application in reasonable times and at reasonable cost. In recent times, this stage involves the development of efficient and easily used computer programs. It is closely comparable to the stage of building equipment in an experimental investigation.

4) Verification
The next step is to test the model against known chemical facts to determine whether the target has been achieved. If quantitative accuracy is being sought, this can be done by various statistical criteria such as the root-mean-square difference between the results of the theoretical model and experimental data. In selecting such a dataset, it is important to make it as broad as possible, while limiting it to experimental facts known to be of high quality. If the results of such a comparison do meet the target requirements, the model may be said to
be validated.

5) Prediction
Finally, if the model has been properly validated according to some such criterion, it may be applied to chemical problems to which the answer is unknown or in dispute. If the experimental dataset is sufficiently broad, there is a reasonable expectation that the results will be accurate to something like the target accuracy. This stage, of course, is the one of
most interest to the larger chemical community.

From - Nobel Lecture: Quantum chemical models by John A. Pople.
Reviews of Modern Physics, Vol. 71, No. 5, October 1999, page 1287.


As a computational scientist, all the stages should be of interest. As a scientific software developer then stages 3 and 4 are of the most interest though often all stages cross our path. Particularly stage 3 where often we as scientists are building our experimental apparatus (computer software) for the experiment ( computer simulation.) Interestingly enough as scientists we often do not build our software to the same standards we build our experimental equipment. Why? I am still trying to formulate the reasons we often seem to lose our scientific training when it comes to scientific software and computation.

Pople On Dirac's Famous Remark

The Schrodinger equation is easily solved for the hydrogen atom and found to give results identical to the earlier treatment of Bohr. With inclusion of the relativistic corrections via the Dirac equation, almost perfect agreement was found with experimental spectroscopic data. However, exact solutions for any other system was not found possible, leading to a famous remark by Dirac in 1929:

"The fundamental laws necessary for the mathematical treatment of a large part of physics and the whole of chemistry are thus completely known, and the difficulty lies only in the fact that application of these laws leads to equations that are too complex to be solved."


This was a cry both of triumph and of despair. It marked the end of the process of fundamental discovery in chemistry but left a colossal mathematical task of implementation. In retrospect, the implied finality of the claim seems excessively bold.


In 1929, there had only been one preliminary approximate quantum-mechanical calculation on the hydrogen molecule by Heitler and London, leading to a value of the bond energy of only about 70% of the experimental value. Nevertheless, the physicists were highly confident and most moved on to study the internal structure of the nucleus during the 1930s. In fact, their boldness was apparently justified, for no significant failure of the full Schrodinger-Dirac treatment has ever been demonstrated.



Nobel Lecture: Quantum chemical models
John A. Pople
Reviews of Modern Physics, Vol. 71, No. 5, October 1999, page 1287.

19 September 2009

Where Is Software Development Really Learned?

Back in April, Bob Martin of ObjectMentor wrote a blog post, entitled “Master Craftsman Teams,” about software development, team structure, and “maybe this is where Agile got it wrong.” It came out on April First and was a good bit of a departure from Agile and XP which Bob has been a major proponent for as long as anyone. Some would even say the post was anti-Agile.

The controversial post generated a lot of reasoned and well thought out responses. It was described by one commenter as, “a thought experiment.” I would tend to agree. Sort of a “let's shake up the status quo” and see what falls out.

Martin's post was very similar in spirit to an essay written by Peter Gill about density functional theory (DFT) called “Obituary: Density Functional Theory (1927-1993).” Peter’s essay style was like… well an obituary, with a bit of tongue in cheek fun, and was intended to generate discussion about DFT and hopefully provide a jump start to physicists and chemists on atomic and molecular theory. Though in his essay, Peter did not propose a new direction as Bob Martin is doing in his post.

Bob proposes, as his title states, a team of craftsman for software development. This team would have at its center a Chief (Master) Programmer along with several Journeyman programmers and several Apprentice programmers per Journeyman. Bob gives the background for this proposal and describes the basic responsibilities within this hierarchy, the cost benefits, and where to find the apprentices for this craftsmanship. Though I disagree with Bob on several of his points, which I will not address now, he has some very interesting points.

In the article Master Craftsman Teams, I feel the real point that really came out of the post and comments was, “Where is software development really learned?” In school …. On-the-job … self-study, or where?

Unfortunately, in the end I think it is primarily a combination of on-the-job and self-study. As for myself, since I have an applied and computational physics background, then I know it was through on-the-job training and self-study. And, except for a few good exceptions, the work environment was not conducive with learning good software development. Many of the experienced scientific programmers just were not interested in improving their skills for software development. They reached a point and said it was good enough or had other issues to on their mind. Thus often many who are learning under them are content not to grow far beyond their mentor.

Now, I naively thought I was just learning the hard way what C.S. students already knew. But the more I interacted with C.S. students the more I realized that often many of them did not learn it in school either if they learned it at all. My guess is they learning it the same way I am learning it except their surroundings may be a better learning environment.

This leads me to believe the sooner a developer gets real hands-on experience developing software the better. Especially with people who care about software development and continually improving at their craft. This is why I am beginning to think internships and mentoring programs within the software industry and concurrently with formal education is really important.

Here at Los Alamos National Laboratory, there are student mentoring programs for high school through graduate school. There can be perfect opportunities here for those who want to do scientific programming. The real problem is hooking up with folks who do not want to rest on their software development laurels but who really want to be better developers and continually grow.

The sooner you start learning and caring about software development the better. And you do not have to wait for school, internships, or a job. There is a lot of information on the internet to start reading about and learning to prepare you for that first internship or job. Just care enough to do it.

Also see Steve McConnell's blog Chief Programmer Team Update that sheds more light on the IBM development in the 60's.

"For the CPT model to work, the Chief Programmer doesn’t have to be 10x as productive as the worst programmer. He has to do the work of eight or nine people put together, which means he has to be 10x as productive as the average programmer, not 10x as productive as the worst. That’s a very tall order" - Steve McConnell