Monday, May 5, 2014

I Learn At edX.org

At least, that's what's written on the T-shirt (er, on the back, where you can't see it on the photo). Indeed, from edX alone I have to my name 7 certificates, another two are in the bag and just waiting to arrive. Number 10 should be Stat2.2x (that you can see on the pic); as for #11, it'll probably be ASTRO1 from ANU (though that might be trumped by the intriguing MAS.S69x and its announced 1-week duration).



Anyway, the great people at edX have been nice enough to send me a T-shirt, so thanks edX!

Sunday, May 4, 2014

Notes from the trenches: Greatest Unsolved Mysteries of the Universe from ANUx

The Australian National University in Canberra have an astrophysics course up on edX called “Greatest Unsolved Mysteries of the Universe”.

I have to admit I only registered when I figured that one of the lecturers − one Brian Schmidt − was a Nobel Prize laureate. I mean, I enjoyed A brief history of time and all, but truth be said, I'm not so much into astrophysics as all that, what with juggling between extremely large quantities (uh, how many megaparsecs did you say?) and extremely small quantities (like, what's an electronvolt?) and generally complex physics. You have to remember, I never really studied physics beyond Newton, Maxwell and Bernoulli. I have a nodding acquaintance with relativity through reading a lot of sci-fi, but in my vocabulary, “quantum” is a passable synonym for “magical”. Still, getting a course from a Nobel prize winner is an opportunity one never really should pass on, if only for bragging rights.

This guy is my teacher. And yeah, he's got a Nobel, sure. Doesn't yours have one too?

Even more so when it's not one course − but four. And not wimpy five-week courses either; no, full-fledged, deep-down, 10-week courses each. A grand total of 40 weeks of tutelage about the finer points of recent astrophysics research, from some of the best-regarded players in the field. So, yeah, I registered for ASTRO1 and ASTRO2 (Exoplanets, due this summer).

So, six weeks into the whole 10-week kaboodle, what to think about it?

First off, it would be unfair to say this is a course by Brian Schmidt. I mean, it is. But it's co-authored and co-hosted by one Paul Francis who should take at least as much credit as Schmidt for the course, and possibly more, if only for his awesome[1] participation on the course discussion forums and his award-worthy dress sense.
I want a waistcoat like this one.
Francis' fast, excitable delivery is also very engaging, but I guess that's subjective. Based on his personal page on ANU's website, he's been quite involved in outreach programmes, and won several awards for science teaching; basically, it shows.

So, what about the course itself? It's brilliant. As expected, I'm struggling a bit. Some notions I seem not to be able to commit to long-term memory. As often, I get the qualitative aspects of the science and find the quantitative bits less involving. Since I tend to put this course on the backburner, I tend to do the homeworks late at night, when I'm not at my most, shall we say, fast-thinking. In spite of this, the course is easy. It's introductory, we're spared the worst calculations (we don't get to do much worse than simple Newtonian point mechanics), all the calculations are (sometimes excessively) spelled out. So I guess if I'm having difficulty with the homework, it's only because I don't put in much more than minimal effort. But that's okay, that's part of the deal. (And anyway, I'm just in a huff because I got one question wrong in 5 whole homework assignments. Fear not; I'm still well on my way to get my 10th edX certificate with this course.)

Generally, each week is devoted to a particular unsolved problem, such as the expansion of the Universe, the nature of the first stars (and why we can't see them), quasars, and so on. Noticeably, the lessons make use of very recent data (some readings from last year only were mentioned in week 6), so it gives a good idea of what astrophysicists actually do nowadays.

One particularly brilliant idea is the weekly mystery, the premise of which is: we've been moved to an alternate, parallel universe, that is superficially like ours (the same laws of physics apply) but quite different in some respects (such as, the night sky is full of brightly-coloured bubbles). Each week we're given some information about the mystery universe; the final will, apparently, test how well we've figured it out. That's fun as it is, but the particular stroke of brilliance is that the information we get is that which we ask for on the discussion forums − that is to say, Paul Francis trawls the forums all week long to find out the most frequently asked-for piece of data, then builds the data set (often complete with figures, “photographs” of the bubbly sky, etc.) and gives us the result as the next week's mystery entry. It creates a great dynamic with the class.

It must also be a dreadful amount of work for the staff; I guess the course isn't going to run with this format every year. All the better reason to stick to ASTRO1.

[1] And that's not a word I use lightly.

Catching up on Stat 2.2x

As I wrote in an earlier post, I regret not having taken Stat2.2x from the start. I have since registered, and it's sort of an uphill battle to get back up to speed. I've viewed the videos up to about the halfpoint of week 2 now, and tried my hand at the problem sets (ungraded of course since the deadline is way past.) I didn't do too badly, but not nearly good enough to pass the overall course.

Let's do the calculations:

  • there are five exercise sets, the lowest score of which is dropped, worth 25% of the grade together (meaning the 25% are split into 4 ES scores). So each ES is worth 6.25% of the overall grade. Three exercise sets are still available to me, so I can get up to 18.75%;
  • last week's midterm was also worth 25% − but that's closed to me now;
  • the final is worth 50% of the grade.

One needs 50% overall to pass the course. A total of 68.75% is still reachable; so a pass is far from impossible. However probabilities is something I never really developped a good intuition for; this takes time. Rushing through the course (doing in three weeks what's meant to be done in five) may not be the best way to do it.
And of course, it means I have to catch up to week 3 today since the corresponding exercise set is due at 1am tonight, Paris time.
But eh, it's a challenge, right?

About the course, then: Stat2.2x follows Stat2.1x, and everything is kind of the same: the lectures are very long and slow, Prof. Adhikari speaks very slowly (with a lovely accent though) and repeats herself quite a lot. That's deliberate: probabilities is one area in which one really should get an intuition for how to approach problems, and going through the material slowly and deliberately is a good way to build that intuition. Still, when you've understood something, it kind of grates to have it repeated three times over fifteen minutes.
The course logistics are a bit different from other MOOCs: all the lectures are released at once, but the exercise sets are released only at the start of the week under scrutiny. The midterm and final have a rather harsh timeframe, as there are only two days between release and due date, and they're on weekends. Since the final alone is worth half of the overall grade, one may have to choose between passing the course and that romantic getaway to Florence one had planned for months before.
In practice, it means one can quickly watch the whole course's material, then review the relevant material just before the exercise sets / midterm / final when they're released. That means going over the material twice, which helps integrating the knowledge, I find.
With a more forgiving schedule for the exams, it would be a perfect arrangement. But then again, left to my own devices, I'd do away with midterms and finals altogether − but that's a topic for another blog post.

Anyway, I'd better stop procrastinating, and learn stuff about hypergeometric probability distributions.

[EDIT: Yay! Made the 6.25% in time.]

Saturday, May 3, 2014

Postmortem: Fundamentals of Immunology part 1 at Rice University

Rice University have put up a Fundamentals of Immunology on edX; part 1 is closing right now. I've been quite successful at it with an overall grade of 90%; it's been rather a lot of work to get to that level so I'm rather proud of it. It's also going to be my second Verified Certificate from edX.


Immunology?

Yeah, you know; the study of the immune system. White blood cells, lymphocytes, the CD4 receptors that VIH binds to, auto-immune reactions (including allergies), the lot.
I've actually got a vested interest in learning about all that, as a lot of people around me have autoimmune diseases; but it's a fascinating subject in its own right. The downside is that, as most medicine-oriented courses (it's based off a pre-med course at Rice) it's heavy on memorization, so be warned, if you take this course or indeed any other course on Immunology, you better be ready to spend hours revising.

What does the course cover?

This is the first part of a two-part course; I expect most of the really hairy stuff to be in the second part.
First off, the course focuses on the operating principles of the immune system in vertebrates, especially humans. Mice and even birds are mentioned at times, but really, it's all about people.

The course starts with an overview of the different types of pathogens (in ascending order of complexity: viruses, bacteria, single-cell eukaryotes such as Giardia, fungi, worms), then a 30'000-feet-high overview of the immune system(s) in higher organisms (plants, fungi, animals) and the distinction between the innate and adaptive immune systems (the latter being specific to vertebrates). Starting with lecture 2, we dive into the details, with an overview of hematopoiesis (how blood cells are made and how they differentiate) and a long list of white blood cell types (myelocytes, lymphocytes, neutrophils, basophils, dendritic cells, B-lymphocytes, T-lymphocytes, etc.) Lecture 3 is a quick run-down of innate immunity.
Most of the rest of the course focuses on B-lymphocytes, the immune system's antibody factories: how antibodies are structured, how B-cells differentiate, how and where they mature, etc. That takes three whole (and information-dense) lectures. The course finishes with a discussion of the complement system, that is to say the molecular process by which pathogens, once identified by antibodies, are neutralized and killed.

So, quite a lot to fit in 6 weeks of lessons. The estimate of 7-10 hours per week on the course presentation page may be a bit higher than what I actually did, but it's not very far off.

Who is the teaching team?

The lecturer is Dr Alma Moon Novotny. (Don't be fooled by her Russian-sounding name, she has a very strong American accent!) She obviously has a long experience of teaching the subject, and makes a lot of effort to make the “memory load”, as she says, lighter. She does it by way of models, cartoons, and analogies. It's a bit strange at first to have cartoonish characters in the slides for a college-level course, but as soon as you realize that you have to learn the essential characteristics of each of these cells by heart, you start thanking her for the fun way in which everything is presented.
In a similar vein, she is generally funny and jokey (for instance calling the stem part of an antibody the “Yoo-hoo! bit” since it's the one that summons other cells) and, well, just fun to listen to. This really helps in such a basically arid subject.
My name is Bond. James "B-cell" Bond.

What about logistics?

The course lasts for 6 weeks, plus two for wrapping up (review, final exam, grading). Each week is generally taken up by one lecture (three lectures are squeezed in the first two weeks), divided in shortish segments of about ten minutes each. Below each segment are one or two ungraded “fact check” questions to make sure you've understood it all.
The lectures are accompanied by two PDF documents: the lecture outline, and the slides themselves. Dr Novotny recommends using the outline to follow along with the lectures; I've been doing a mix of reviewing the outline before watching the lectures, and following along with the slides. The outlines and slides are a great help for reviewing; to prepare for the quizzes and final exam I eventually printed them all out and carried them around everywhere.
(Phew, revising lessons on the bus: hadn't happened to me in fifteen years!)
A quiz wraps up each week. Somewhat unusually for MOOCs, the quizzes are “closed-book”, which is to say you're not supposed to have the course material (or indeed anything else) at hand while taking them. There's no way to enforce the policy though, so it's all a matter of honour on the students' side. (To be perfectly honest, I hadn't understood they were closed-book until the third quiz. Note however that I didn't actually get much better grades on the first two, when I had the outlines etc. at hand, than on the other four or indeed the final exam, for which I did adhere to the closed-book policy).

The course is wrapped-up with a longish 60-question final exam covering the whole course, also closed-book.

One question per page… doesn't quite mesh with the edX navigation system

As usual for MOOCs, there is no textbook (which is why the outlines are so detailed). Dr Novotny does provide a handful of links to interesting resources on the Internet though; while revising, I found (as she mentions) that the relevant Wikipedia pages are actually very good.

My impressions

I enjoyed the course a lot. First because I learned a lot of stuff, then because Dr Novotny is simply a joy to listen to.
The less enjoyable parts were, of course, the quizzes. I doubt there's anybody on Earth who actually likes doing quizzes… It didn't help that some of them were mis-coded (this was obviously the first run of the course and the staff obviously had to get to grips with the edX platform; they were, however, very responsive whenever errors were flagged on the forums). We're evidently far from the very sophisticated 7.00x Introduction to Biology from MIT with its wealth of interactive tools instead of simple yes/no/maybe quizzes. However, it's also obvious the Rice team hardly had the same budget as the MIT one's for producing the course − and it would be unfair to decry the course for not being up to the very best course I've ever seen. This Immunology course is all that can be reasonably expected, and more.
Dr Novotny with an antibody

A quick note: as I said in the introduction, I paid for the Verified certificate. Not so much because I think the certificate will be helpful in my career (I don't see how it would) but because it's cheap (25 USD, about 20 euros), it's a way to indicate appreciation for the work being done, and it's an added motivator: having paid, I'm less likely to drop the course, even if it's hard work.

Overall, I'm eagerly anticipating part 2 (where we'll learn all about T lymphocytes). Do be warned though, if you want to take up this course: it's a lot of work, and a lot of it is unfortunately (but unavoidably) about memorizing stuff.

[Edit] And now the certificate's arrived!


Friday, May 2, 2014

Postmortem: Bioinformatics at Peking University

The Bioinformatics course at Beijing (or is that Peking? I never know) University is over. It's still being graded, but having had 10/10 for each homework and 97% on the final, it's not an overly wild guess to say it'll be my first Coursera certificate (or “Statement of Achievement” in Courserese).

So, what is Bioinformatics?

Bioinformatics is “the application of computer science to solve biological problems”. To say it's a growing field would be an understatement: the growth of biological data has been exponential, necessitating innovative data management and data analysis techniques; it's fair to say that nowadays, in a lot of fields of biological research, more work is being done with computers than with, say, Petri dishes or lab mice. Taken from the opposite angle, biological research labs are at the forefront of the “big data” revolution. Lists of innovative big data companies include institutions such as Mount Sinai's Icahn School of Medicine (they are also well-known as a big MongoDB customer, if I'm not mistaken).
Rigorously, “bioinformatics” comprises an understanding of the type of problems bioinformaticians face, and the algorithms they use to solve them. By extension, it also includes the major databases of publicly-available bioinformatics data on the Internet − and truth be told, I didn't think there were so many of them!
As is usual in fields associated with academic research, bioinformatics is a field in which open source is prevalent, both in terms of the software itself (indeed, algorithms and tools that are not properly described in a peer-reviewed paper have little chance of being widely used) and in terms of the actual data. To repeat myself: the amount of data in freely-available databases hosted by such organizations as the National Center for Biotechnology Information or the European Bioinformatics Institute is staggering. Theoretically, any private individual with an Internet connection, or indeed company, could do effective biological research in silico; I'm guessing we're just at the beginning of a wave of bioinformatics startups following the lead of such as 23andme. It'll be interesting to see.

What does the course cover?

The Peking course covers quite a lot, from a rather comprehensive (and welcome) review of the history of bioinformatics, the present situation of the field, and what we can expect in the near future, to alternating descriptions of the major algorithms and/or techniques in use in the field and exploration of the available resources. It ends with a couple of case studies highlighting how the various techniques and databases surveyed in the course were integrated by actual researchers to investigate real-world issues.

Who is the teaching team?

There are two main professors: Drs Liping Wei and Ge Gao. Dr Wei generally handles the high-level stuff (e.g. background information, overviews of databases, etc.) while Dr Gao dives into the details of algorithms. Both are evidently highly qualified; the course draws tightly on their own research and contributions, which is very welcome as it makes the course so much more concrete.

What about logistics?

The course lasts for 6 weeks; two topics are covered each week, with a series of video lectures and a quiz; except for the last week, when the lectures are about case studies and the quizzes are replaced by a 40-question final exam.
It is notable that the course is bilingual; that is to say, it is offered simultaneously in Chinese and English. The Chinese lectures have slides with classroom shots inserted (meaning you actually see the professor speaking); the English lectures are slides-only.
In addition to the lectures there are supplementary videos such as lab visits and student presentations. These tend to be Chinese-only, so I skipped over most. There are also a couple of interviews of famous bioinformaticians, but I didn't find these so interesting.

My impressions

It's undeniable that the team take pains to be welcoming. It's also undeniable that the content is actually of a pretty good level. But there are a couple of problematic aspects with this course.
The first is − and it's horrible to say this, not being a native English speaker myself − that the speakers' English is not so great. It's not so much that they make mistakes, but rather, their intonation and rhythm is… well, they're obviously reading from a transcript. Dr Gao's voice droning formulae (ecks-aye-jay-plus-ecks-jay-plus-one-aye-equals-ecks-aye-plus-one-jay) for long minutes was almost enough to take me out of the course altogether. Thankfully, the subject matter is interesting, so I could stick to it with a little effort. In the end, I viewed the lectures at 1.5x speed to compensate for the speaker's slow diction, and I referred back to the transcripts when I had doubts.
The second problem is perhaps due to Coursera's platform: the quizzes are, well, just quizzes. It feels strange, and generally wrong, to have an algorithmics course in which you do no coding at all. Generally speaking, you have three tries to answer questions which mostly have three or four options… At one point I was so immensely tired I almost dropped the course; instead I deprioritized it and spent a minimal amount of time on it. So, while I did learn a bunch of stuff about bioinformatics, I can hardly say that my final grade (which will be in the high nineties) reflects my mastery of the subject.
Strangely enough, the final exam was possibly the most interesting part of the course, as some questions required us to go search for information by ourselves (which means first identifying the right database to query, finding how to query it, etc.) Possibly, a vast improvement of the course would be to scrap about half the quizzes and replace them by practical case studies: give students a set of data (gene names, diseases, etc.) and send them off to search for information using the tools discussed in the lectures. As they stand, the quizzes help little in teaching.

Do not, however, let that discourage you. If you're interested in learning about bioinformatics, certainly there could be much worse options than this course. It may not make you a bioinformatics researcher, but it will at least give you a handle on where to get started.

Just started: M202 MongoDB Advanced Deployment and Operations

So MongoDB has a new course.

Why take this MOOC?

Because M102 felt a bit too basic at times. More generally, that's one of the few courses I'm taking that have an immediate bearing on my career (which is why I've dealt with my bosses to spend a few hours a week doing the course, on company time); I believe NoSQL databases (for good or for worse − I'm hardly a fanatic either way) are here to stay and do solve, or help solving, a number of real problems. Taking this course is, I think, the surest way to get a leg up on the most popular NoSQL database around.

Why do MongoDB offer the course anyway?

Generally, open source software companies earn money through selling a mix of training, support, and consulting / professional services. So creating high-quality training content and putting it on the web for free is, well, rather unusual.
That's because MongoDB's business model seems more focused on selling cloud-based additional services (such as online backup) as well as the usual support and consulting. It's a gamble − on the one hand, it deprives them of lucrative training business; on the other hand, the more people know the product (and the additional services that MongoDB sells) the more the product is likely to be used and the more additional services are likely to be bought.
Time will tell if it's a winning gamble; for my part I applaud heartily this initiative.

Fifo goes to eleven


In spite of an eventful week (including an unexpected couple of days in hospital; nothing serious though), my progress bar has just reached 58% on the Analytics course at MIT, where 55% is deemed a pass. So… that'll be my 11th MOOC certificate (9th from edX alone)!
I'd uncork the champagne, only I haven't got any at home (and anyway alcohol doesn't mix very well with the painkillers).