Thursday, April 20, 2023

Messy Miller Experiments

Research into the chemistry of the origin of life got a kickstart with the 1953 publication by Stanley Miller reporting on his experiments that produced amino acids, the building blocks of proteins, from simple substances (H2, CH4, NH3, H2O). The energy source used was a spark discharge (emulating lightning) so these are sometimes referred to as spark discharge experiments. They are also known as Urey-Miller or Miller-Urey experiments, attaching the name of Miller’s supervisor, Harold Urey. While Urey had initially tried to discourage Miller from origin-of-life chemistry, thinking it would not yield anything useful, he was very supportive by giving Miller full credit as sole author on the now-famous 1953 paper.

 

Seventy years later, such experiments continue. We’ve learned a lot since then. One of the important lessons is that to get a wealth of organic compounds in your experiments, you have to embrace the messy! A recent publication by Root-Bernstein and colleagues is the latest iteration. Here’s a picture of the abstract. (The article is open-access so you can read it in full; DOI: 10.3390/life13020265)

 


As the title suggests, the experiments produce a rich mess: not just amino acids but also sugars, nucleobases, lipids, and even oligomers – the linking of building blocks to make larger molecules that life utilizes. To the solution they add a particular sea salt from the Mediterranean along with hydroxyapatite (a phosphate salt) and magnesium sulfate (Epsom salt), common salts you might expect to be present in prebiotic chemistry. They ran pre-tests to make sure no ATP or other “biological” molecules were present. Other than that, they tried to stick to Miller’s original initial conditions. One crucial way in which the protocol differed from the original, is that instead of a one-pass reaction (Miller’s was seven days), the researchers regularly reintroduced the “atmospheric gases” (H2, CH4, NH3) every seven days and ran the experiment for multiple weeks. This allows fresh “nutrients” to enter the system so that the chemistry can progress much further – and get messier!

 

At regular intervals, they remove some sample solution to figure out what’s in it. As expected, within the first week, amino acids show up similar to the ones Miller found. Sugars also show up early. This is expected since the Strecker reaction for producing amino acids relies on the building block formaldehyde, which is also the building block for producing sugars in the formose reaction. By the end of the second week, nucleotides are found, i.e., nucleobases were synthesized. Again, not surprising, since precursors such as HCN and formamide are expected to play a role. The next few weeks bring fatty acids and peptides (oligomers of amino acids) into the mix. There are a couple of dicarboxylic acids (malonic and succinic) and even a few sterols. It’s a very impressive mix – you’ve got stuff for membranes, proteins, sugars, nucleic acids, and metabolic diacids.

 

The main question that arises in this embarrassment of riches is whether there might be contamination, especially when “regassing” or when removing samples for testing. The authors have expected this question and discussing how they have been careful in their protocols, and how they have tested for errant contamination. While contamination cannot be completely ruled out, it seems to me that they have done due diligence and there don’t seem to be any telltale biomarkers present that you’d find in contaminated samples. The other thing that jumped out at me was the presence of the (sulfur-containing) amino acid cysteine. Cysteine is sort of a magic amino acid that catalyzes all manner of reactions. You can produce it by adding H2S to a spark discharge experiment, but in this instance, H2S was probably produced in situ by reduction of sulfate. This means that thioesters could come into play, although no mention is made of them in the paper. I didn’t expect thioesters to be directly observed because they would hydrolyze easily under these experimental conditions.

 

The authors rightly emphasize the messiness and the advantages of what they call “dirty” experiments. I’m in agreement. The analysis is painful, and teasing apart what’s going on in such reaction mixtures will continue to be challenging, but I think this is the right approach to make progress in the field. Hurrah for the mess!

Saturday, April 15, 2023

Magic of Babel

The biblical tower of Babel fell into ruin because of confusion among people who spoke different languages. Presumably it was a severe breakdown in communication, although the Bible doesn’t provide details. Building on this idea, the author R. F. Kuang has written a novel, aptly named Babel, with a magic system that exploits the differences of meaning between different languages. What is lost in translation becomes a source of power. It’s a difference that makes the difference!

 


While Babel has been compared to Jonathan Strange & Mr Norrell, other than the setting of the early nineteenth century and the common theme of magic, the two novels are very different. Babel is another iteration of the coming-of-age story whereby a young protagonist is initiated into the study of magic while making new friends – we’ve seen this before many times, it’s a very effective trope. While this is interesting in the first half of the book, the second half focuses on race, colonialism, power, economics, and the lengths people will go through to maintain the hegemony of the status quo. As someone from a Commonwealth country, I’m familiar with the historical setting of Pax Britannica, and I also had the experience as a multilingual international student of being in a foreign place with no other countrymen (at my college). That being said, I feel that Babel was overall depressing, and I think the sociological arc of the story is weaker – it feels like the author is trying too hard to push a particular point of view.

 

The tower of Babel is the centerpiece of Oxford because that’s where the magic happens. The ascendance and economic might of the British Empire and its East India Company maintain hegemony because they fashion magical technologies by etching words in different languages on silver. When someone fluent in those languages speaks the words, it imbues the silver object with magic that enhances a particular technology. For example, it can make cannons fire more truly, carriages travel faster or carry heavy loads, and caravels to more ably weather storms at sea. Metal-etching and language as the foundation of magic isn’t novel – the recent Foundryside series employs a related system – but the system in Babel is more interesting. Kuang’s background in multiple languages enhances her novel, and in my opinion, these are the strongest and most interesting parts.

 

When our main protagonist Robin and his fellow first-year students get their first tour of Babel, a professor says: “Now that you’re part of the tower… the tower knows you.” They’ve just given their blood so that they will be allowed safe entry and exit into the hallowed halls of Babel. Gotta keep the riffraff out. (Foundryside has a similar situation.) Then the professor intones: “You’re in the place where magic is made. It’s got all the trappings of a modern university, but at its heart, Babel isn’t so different from the alchemists’ lairs of old. But unlike the alchemists, we’ve actually figured out the key to the transformation of a thing. It’s not in the material substance. It’s in the name.”

 

I like that Kuang makes the connection to alchemy and the transformation of substances (transmutation), but then emphasizes that it’s not so much its ‘elementary’ material as its name that makes the difference. The power of names and naming is spread across myriad ancient cultures, and still holds significance today. It’s unfortunate that our hypermodern world has too a large extent lost this significance. I also think the idea that information underlies matter, rather than the other way around, is particularly intriguing – and there might be evidence deep in fundamental physics – but that’s a story for another time.

 

In the introductory class to Translation Theory, the professor says: “The first lesson any good translator internalizes is that there exists no one-to-one correlation between words or even concepts from one language to another… Language does not exist as a nomenclature for a set of universal concepts… How do we render the French esprit into English?” This then leads to the “three principles of translation”: A student recites these: “First, that the translation conveys a complete and accurate idea of the original. Second, that the translation mirrors the style and manner of writing of the original. And third, that the translation should read with all the ease of the original composition.”

 

The emphasis of really understanding more than one language deeply and fluidly moving between one language and another is magical, even in our non-magical world. I have first-hand experience trying to learn a new language that I did not grow up speaking. It’s hard work! I haven’t achieved depth or fluency, but I know it’s possible. And the Babel students work insanely hard at translation and getting to the tangled evolutionary roots of words in different languages. The professor also hints at a primordial ‘Adamic’ language that may have been lost while making hints to biblical Babel. A similar concept in Foundryside takes centre stage, while it’s merely a curiosity in Babel.

 

An example of how the magic works by exploiting the differences of word meanings in languages is provided by the professor: “The Greek word idiotes can mean a fool, as our idiot implies. But it also carries the definition of one who is private, unengaged with worldly affairs – his idiocy is derived not from lack of natural faculties, but from ignorance and lack of education. When we translate idiotes to idiot, it has the effect of removing knowledge. This [silver] bar, then, can make you forget, quite abruptly things you thought you’d learned. Very nice when you’re trying to get enemy spies to forget what they’ve seen.”

 

But there are problems. Silver tarnishes over time and must continue to be ‘maintained’ by cleaning the dross and re-inscribing. The professor explains that depending on the desired effect, you might need larger quantities of silver or higher purity. Most of the bars are alloys, and the continual need for repair and maintenance keeps the scholars of Babel rolling in riches. They have a monopoly. The other interesting aspect comes up in a conversation among students. As languages evolve and start to share more similarities, the erasure of differences causes the silver bars to lose power. This anchors an aspect of the story where Babel is interested in Oriental languages and the Far East. The larger differences in translation result in more ‘raw’ magical power to be exploited. This gives the book an international cast and emphasizes a variety of languages. The interplay between language and magic in Babel is deeper and more interesting than in Foundryside, and in my opinion, it’s the most magical part of the novel.

Thursday, April 13, 2023

Reminiscing Catan

Klaus Teuber passed away last week. For those who don’t recognize his name, Teuber designed the boardgame Settlers of Catan. It was one of two games that I discovered in the ‘90s that led to my renewed interest in boardgames. (The other was Richard Garfield’s RoboRally). Settlers of Catan led to a renaissance in boardgames that hit a Goldilocks sweet spot. Not too long and complicated. Not too short and trivial. It has a nice balance of luck and strategy, and there was both a competitive and cooperative aspect to winning gameplay.

 

In those early years, I played hundreds of games of Settlers, as we called it then. (Now the game is just called Catan.) I still have my first (English) edition. Although the box cover is slightly beat-up, the components are still in good condition. But with the arrival of newer games, Settlers began to lose its shine. I will still happily play it with the right crowd, i.e., one that doesn’t devolve into analysis-paralysis or hyper-competitiveness. For me, the sweet spot is a four-player game that takes 45-60 minutes. It’s much better as a relaxing, low-key, game of building and trading. And occasionally frustrating your opponents.

 

The first edition does not mesh with any of the expansions, so I’ve never bought them. Also, the expansions add to playing time. I like the Seafarers expansion. It adds 15-30 minutes, depending on the scenario, but provides an element of discovery (which is fun!) and also mitigates the sheep problem. (If you’ve played Settlers a lot, you know what I mean.) On the other hand, the Cities and Knights expansion easily adds an hour or more, and results in more hindrance strategies. I tired of it after my first game, although I did give it a chance by playing it more than once. There are now many themed versions of Settlers of Catan. I only own one: Settlers of the Stone Age because I like the history-evolution theme to it. The gameplay is, overall, slightly inferior to the original, but I like the theme enough to play it sporadically.

 

I looked up the list of games designed by Teuber to see how many I’ve played. Elasund and Entdecker are games I used to own but gave away to friends as I accumulated too many games and they were hardly played by me. Teuber hasn’t actually designed that many games outside of the Catan spinoffs and expansions. Two famous ones that I haven’t played are Barbarossa and Hoity-Toity. Looking at my own collection, the designer whose games I have collected the most is Reiner Knizia, a prolific game designer who has elegant designs likely influenced by his background in mathematics.

 

Back in the day, these “famous” German game designers worked on games as a side-project. They had full-time jobs. Teuber worked as a dental technician. Knizia worked in finance. Both finally quit those jobs to become full-time game designers in the late ‘90s. What is it about the Germans and their side hobbies? One of the most influential theories of the origin of life came from a patent lawyer, Gunter Wachtershauser, who quietly toiled away from years before unveiling his metabolism-based iron-sulfur world. It shook up the origin-of-life community and opened new lines of research.

 

Nowadays, my game-playing has ebbed to a low level. The pandemic didn’t help – although there was a resurgence in my playing the boardgame Pandemic during that time. I read more now; I suppose that’s my new hobby. But I’m always up for a good game. As long as it doesn’t take too long, and has an appropriate mix of luck and strategy commensurate with the time I invest in playing it. Settlers of Catan paved the way for the myriad choices available today, thanks to Klaus Teuber.

Wednesday, April 5, 2023

Stochastic Parrots

Digging into ChatGPT and Large Language Models (LLMs) is leading me down a rabbit-hole. In my last post, I looked at a paper discussing whether such models ‘understand’ language, and if so, how different might it be from how humans learn. The crux was whether statistical correlations are sufficient mimics that can, for practical considerations, substitute for knowing causal mechanisms. We don’t really understand what these LLMs are doing when they seemingly come up with novel and surprising responses, which they have not been ‘trained’ for (kinda, sorta). It’s a black box even to those who developed such AIs.

 

Today, I’m looking at a different paper with a catchy title: “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?” The article is open-access at the Association for Computing Machinery digital library linked here. It’s a little more technical, and assumes some familiarity with LLMs. Since this article is from 2021, it does not include the latest version of ChatGPT or some of the newer-released chatbots. But they have more than enough data to make the key arguments in their paper. The background information provided in the paper was very useful for a novice like me to get a brief history of developing such models and a brief overview of the different LLMs out there. There’s a bunch of them!

 

Whenever you query ChatGPT, it takes up energy and compute time. How much energy? There are various estimates on the internet and the kWh per query number may look small, but a quick calculation will show that it is 2-3 orders of magnitude (i.e., 100-1000 times) more than the human brain. LLMs are energy guzzlers compared to us. This doesn’t account for the amount of energy used to train the LLMs in the first place to get them to the user-friendly stage. The bigger the LLM, the more energy it guzzles.

 

Reading about the training data was an eye-opener. I’m a computational chemist so I’m familiar with using data to train computational models. An important thing to keep in mind is GIGO: Garbage In, Garbage Out. How well your program will perform its specific task depends on the quality of your training data. Where does the training data come from for something like ChatGPT? It’s not just the things you might expect like Wikipedia, digitized books, and vetted data repositories. Turns out there are very large datasets known as Common Crawl. Sources include “scraping outbound links from Reddit” and other user-generated content such as Twitter. There are biases in those data sets in terms of who the users are and what they might discuss on the internet. We shouldn’t be surprised at the misogyny that shows up from chatbots. The article authors provide numerous examples of why these issues crop up. I recommend reading their article in full.

 

Given that the amount of garbage on the internet is increasing exponentially faster than anything else, the worse things get. Over the years, I’ve noticed more wrong things (related to chemical data) show up in the top few hits on a standard Google search. Search has in fact become more tedious when I’m looking for good data on specific things I care about. For those researchers who work in Natural Language Processing and are trying to move into Natural Language Understanding, the approach of just throwing more scraped data is likely going to make things harder especially when resources are being thrown at these LLM approaches. Is Big Data the answer? The authors make the point that “[language] coherence is in the eye of the beholder… human communication relies on the interpretation of implicit meaning conveyed between individuals… [it] is a jointly constructed activity… even when we don’t know the person who generated the language we are interpreting, we… [intuit] what common ground we think they share with us, and use this in interpreting their words.”

 

The authors refer to LLMs as Stochastic Parrots. This is an apt name given what these models are doing by applying statistics and Bayesian probabilities to constructing plausible-sounding text. We should think carefully about taking advice from these parrots, and certainly double-check what they’re telling us. But we’re lazy. Why do the hard work that you can outsource to a machine? Isn’t it just a useful tool? Like a calculator? Well, partly yes, but every year I see students punch their calculators and write nonsense answers. You could call it user error, but it’s an error often associated with not understanding how the calculator parses the input you provide to generate the output. I could make an analogy to the workings of a chatbot. Except that things are even fuzzier. There’s now a stochastic wrench thrown into the works. But all that being said, I’m still finding the stochastic parrot interesting to interact with – it gives me a sense of what an alien intelligence might be like.

Tuesday, April 4, 2023

Alien Understanding

As I’ve been playing with ChatGPT, I’ve also been reading about Large Language Model AIs. It’s clear to me that ChatGPT doesn’t understand chemistry in the same way humans do, nor should I expect it to given what it is programmed to output; but does it understand language? I suppose that depends on what one means by “understand” and what one means by “language”. The distinction that comes to my mind is between syntax and semantics. My opinion is that ChatGPT mimics a human speaker because it is well-trained on syntax. We humans feel that it is an effective mimic because in a conversation, we have evolved to impute semantic information to any conversation partners, human, animal, or machine.

 

A useful primer that’s accessible to the non-expert is a recent short four-page article by Melanie Mitchell and David Krakauer. (“The debate over understanding in AI’s large language models”, DOI: 10.1073/pnas.2215907120.) Here are some of their statements that I felt clarified the key questions or highlighted the key issues.

 

“… the current debate suggests a fascinating divergence in how to think about understanding in intelligent systems, in particular the contrast between mental models that rely on statistical correlations and those that rely on causal mechanisms.”

 

“[In the past], the oft-noted brittleness of these AI systems – their unpredictable errors and lack of robust generalization abilities – are key indicators of their lack of understanding… However, [present LLMs]… can produce astonishingly humanlike text, conversation, and in some cases, what seems like human reasoning abilities, even though the models were not explicitly trained to reason.”

 

“… such networks improve significantly as their number of parameters and size of training corpora are scaled up… [some claim] will lead to human-level intelligence and understanding, given sufficiently large networks and training datasets.”

 

“…[others argue that LLMs] cannot possess understanding because they have no experience of mental models of the world; their training in predicting words in vast collections of text has taught them the form of language but not the meaning.”

 

“… the Eliza effect named after the 1960s chatbot… refers to our human tendency to attribute understanding and agency to machines with even the faintest hint of humanlike language or behavior.”

 

The crux is that we don’t exactly know, down to a nuts-and-bolts level (or should that be bits-and-bytes), what human understanding is. We have mental models – conceptual, abstract, and tied together by particular notions of causality – and we perhaps assume that such models aren’t solely statistical correlations. In this model, we have a self. And this self needs to function in the real world of blood-and-guts and nuts-and-bolts. As we interact with our environment, which may include other entities like ourselves, we use what we have learned to make predictions of what might happen next. We have expectations based on our conceptual models, and we are surprised when reality doesn’t meet expectation in a particular functional test. Then we update our model.

 

One argument against LLMs having human-like understanding is that they “exhibit extraordinary formal linguistic competence… they still lack the conceptual understanding needed for humanlike functional language… and use language in the real world. An interesting parallel can be made between this kind of functional understanding and the success of formal mathematical techniques applied to physical theories. For example, a long-standing criticism of quantum mechanics is that it provides an effective means of calculation without providing conceptual understanding.” As someone who teaches quantum mechanics, this analogy strikes home for me. But maybe it’s our present lack of understanding of our own understanding that’s the roadblock. And if we can learn how LLMs “understand” in their own way, that might shed more light on what human understanding means. It’s unclear that the lack of embodiment in present LLMs constitutes a significant hurdle, but one could conceive outfitting a physically mobile robot with an LLM. Such an integration could prove interesting. And scary.

 

If future improved LLMs cannot create something that resembles a rich-enough mental model that seems akin to humans, can they at least do what is functionally equivalent? Mitchell and Krakauer suggest that when we see “large systems of statistical correlations produce abilities that are functionally equivalent to human understanding… [or] enable new forms of higher-order logic that humans are incapable of accessing… will it still make sense to call such correlations ‘spurious’ or the resulting solutions ‘shortcuts’? And would it make sense to see the systems’ behavior not as ‘competence without comprehension’ but as a new, nonhuman form of understanding?” They also caution us not to jump to conclusions when trying to interpret how LLMs perform on our “tests” of comprehension based on our own assumptions of how human understanding works – something we don’t really understand at a deep level. It’s perhaps an alien understanding. We might anthropomorphize such alien understanding in sci-fi literature and media, but perhaps that’s the only way we as humans can try to understanding something beyond our ken. If anything, all this reminds me that we need to remain humble about what we think we know. We’re limited by our own human embodied understanding.

Thursday, March 30, 2023

Chatbot Study Help

While ChatGPT doesn’t actually understand or know chemistry the way humans do, can it still help students in some way? Since my first chat, I’ve come up with two possibilities to leverage the generative capabilities of GPT: study guides and test questions. These are two things that students regularly ask for. So, if a professor is unwilling to provide them, maybe ChatGPT will oblige. (I have been revamping my G-chem study guides this semester.)

 

I asked GPT: “Can you help me put together a study guide for my General Chemistry 1 class? What should I know or know how to do to prepare for the final exam?” It came up with a bulleted list that mirrored much of what we cover in G-Chem 1 (atomic structure, chemical bonding, chemical reactions, gas laws, periodic trends) although it added some G-Chem 2 topics as well. Here’s one of the items:

 

GPT: Periodic Trends.

·      Atomic Radius

·      Ionization Energy

·      Electron Affinity

·      Electronegativity

 

And it added a nice study tip:

 

GPT: In addition to understanding these concepts, make sure to practice solving problems related to each topic. You can find practice problems in your textbook, online, or from your professor. Good luck with your studies!

 

To drill down, I asked: “Under Periodic Trends, could you be more specific as to what I should be able to do or explain.” It does a nice job. Here’s the first item as an example:

 

GPT: Atomic Radius

·      Understand what atomic radius is and how it changes across a period and down a group.

·      Be able to explain the factors that influence atomic radius (such as the number of electron shells, effective nuclear charge, and shielding effect).

·      Be able to predict which element has a larger atomic radius in a given pair of elements.

 

I asked GPT to provide me a sample question that covers the three things on the list for Atomic Radius. (I did not choose Ionization Energy or Electron Affinity because they are trickier.)

 

GPT: Explain how atomic radius changes across a period and down a group, and identify the factors that influence atomic radius. Using this information, predict which element has a larger atomic radius in each of the following pairs:

a.     Cl or S

b.     Ca or Sr

c.      O or F

Here’s my partly erroneous response: “As the number of protons increases, it pulls the valence electrons closer to it. For this reason, atomic radius should decrease going across a period since the number of protons increases. For the same reason, atomic radius should decrease going down a group because the number of protons increases. Applying it to the pairs: (a) Cl will be smaller than S, (b) Sr will be smaller than Ba, and (c) O will be smaller than F.”

 

GPT encourages me (Great effort!) and agrees with my initial response. Then it tells me I got (a) and (c) correct and (b) wrong.

 

GPT: For part (b), your answer is not correct. Sr has a larger atomic radius than Ca. This is because Sr is located below Ca in the same group, and as we move down a group, the atomic radius increases due to the addition of new energy levels.

 

I ask GPT to help explain what adding new energy levels means, and I ask more clarifying questions when it brings up new terminology (such as shielding). It does a decent job in its responses. Then things go awry. I ask it to help me use shielding in a better response for (b).

 

GPT: to improve your answer to (b), you could say that Sr has a larger atomic radius than Ba because, even though they both have the same number of energy levels, the valence electrons in Sr experience less effective nuclear charge than the valence electrons in Ba due to increased shielding. Therefore, Sr has a larger atomic radius than Ba.

 

Eeeeks! The original question was Sr vs Ca. Now Ba gets thrown in, and the answer is wrong. I say I’m confused because Ba is below Sr on the periodic table. GPT apologizes and corrects itself. We do some back and forth on when to use shielding in an argument and it gets things half right and half wrong.

 

Then I go back to an earlier problem and say: “Earlier you said I was correct when I said O was smaller than F, but now I think I might be wrong. Was I wrong earlier?” GPT assures me that I was right but spouts a nonsensical explanation. After some back and forth it corrects itself apologetically. We then go through some examples that include cations and anions, and I essentially see the same thing half-right, half-wrong explanations.

 

In a new chat, I decided to try a different tact on a different topic: “I need practice with stoichiometry in my chem class. Exam coming up. Can you help generate a practice question for me? (Just the question no answer please.)”

 

GPT: If 5.00 grams of calcium reacts completely with excess hydrochloric acid, how many grams of calcium chloride will be produced?

 

That’s a good start. I ask for a hint to help me get started, and it obliges with an appropriate and helpful hint. I then claim to be stuck: “Okay, I'm trying to write the equation: Ca + HCl --> CaCl2 but where does the H go?” GPT responds correctly that H2 gas is released as a byproduct and provides a correctly balanced chemical equation. I then provide a wrong answer. Then, impressively, it tells me the detailed correct steps I should follow to do the calculation. And even better at the end it says…

 

GPT: “So, without giving away the full answer, I would suggest you to follow these steps to solve the problem. Let me know if you have any further questions!”

 

It might be getting the hang of this hint-providing approach. I tell it another wrong answer and things go awry when it tells me in an encouraging tone that I’m correct but suggests I check my sig figs. I say “Oh wait, I made a calculator error” and proceed to give it the correct answer. GPT confuses itself, probably because I used the word “error” and tells me I’m wrong. It then methodically works out the problem and gets the same answer I did.

 

I follow-up asking for a harder question that requires me to “know limiting reactant and %yield”. GPT obliges. Good question. Wrong answer because it picks the wrong limiting reactant, misled by me. When the resulting yield is over 100%, it suggests that there may be error or inefficiency in the reaction. I ask for some examples and it provides a plausible list. I try asking for a harder question with different reaction stoichiometry and GPT chokes on balancing chemical reactions.

 

In a new chat, I try a different topic: “I have a quiz coming up on intermolecular forces in chem. Could you tell me what are the main things I need to know?” GPT generates the key points and does a very nice job at the G-Chem level even though I didn’t mention General Chemistry in the prompt. I ask it to generate three different questions to test different aspects of intermolecular forces. The questions are reasonable although two are essentially multiple choice. Not surprisingly, it can defend the correct answers to the multiple choice questions well. And it goes awry with the question that requires explanation. At one point it makes a claim that correlates hydrogen bonds to atom size when I ask follow-up questions about the explanations. I ask it for harder questions. Once again, reasonable questions, but if pressed for answers, GPT spouts partial nonsense.

 

When I ask GPT to help me make the main points of intermolecular forces memorable with a poem, it obliges in a whimsical ditty that’s better than anything I could come up with. The poem has meter and rhyme. But it doesn’t help with knowing the relative strengths of the forces. So I ask for another poem that incorporates these, and it obliges with five stanzas – the fifth being a summary. I claim the poem is too long and ask it for a “memonic [sic] or memory aid”. Here’s the response:

 

GPT: Dancing Elephants Don't Hug Intensely. Each letter in the phrase represents one type of intermolecular force, in order from weakest to strongest:

·      Dispersion forces

·      Dipole-dipole interactions

·      Hydrogen bonding

·      Ion-dipole interactions

 

I’m not sure where the Elephants come from but the image in my mind is stark. Hard to unforget! Not sure if I’ll remember the mnemonic.

 

Here’s my summary thus far: GPT can generate reasonable study guides. By asking it for more details and “what should I be able to do?”, one can get a pretty good comprehensive guide. GPT can generate a variety of test questions. By specifically providing categories to explicitly include complexity into the question, one can generate decent mock exam questions. I tell my students this is a good skill when preparing for an exam, and that doing so in groups is even better. I could conceive of an exercise that gets students to think about what should go into a good exam question and have them use GPT as an aid. The study guide points can be fed back into the test question generation process.

 

But asking GPT for help answering questions is where things become problematic. It provides plausible sounding answers but there’s no check when nonsense is generated. I think it’s useful for students to see this too. I find it interesting to read GPT’s responses because sometimes it reminds me of a student who crams information without understanding and then does a data dump. Except GPT is much smoother sounding than the student. Perhaps it can lead to a helpful discussion on what they should look out for when they use GPT, and what it means to actually understand something. Maybe it will help students be more reflective and ask themselves if they really know something or if their knowledge is just plausible-GPT-like. I regularly tell students that they should speak or write out explanations, and that the act of doing so helps with self-clarification. If seeing GPT in action helps them with such metacognitive awareness, that’s probably a good thing.

Wednesday, March 29, 2023

Study Guide Feedback

Last week, which marked the middle of the semester, I asked my G-Chem 2 class for feedback on my new study guides. I’m teaching the Honors section this semester so it’s a small class. The averages on the two midterms are the highest I’ve had, and my exams have not gotten any easier (nor harder), in my opinion. So it may be that the students are being helped tremendously by the new study guide format. But more likely, they are being helped a bit, and they are already very strong students. Furthermore, size-wise this is one of the smallest classes I’ve taught. The students are engaged and ask good clarifying questions in class. All this to say that a number of factors are contributing to the strong student performance, and it’s not just the study guides.

 

Here are the four questions I asked, a summary of the students’ feedback, and some of my thoughts in response.

 

Q1: How have you been using the study guide? (Be detailed and honest, please!)

 

Some students are using this the way I intended after each class. Others are only using them as an exam crops up. One student mentioned going through them once a week (covering three class periods each time). Since the responses are anonymous, I don’t know if there’s any correlation between performance on exams and how the study guides are used. Two students explicitly mentioned using them to prepare for the pop-quizzes. One student mentioned working with others on it to clarify different ideas and answers. One splits up the conceptual questions from the “test yourself” problems, looks at the first group more regularly and then does the problems close to the exam as sort of a practice-test. All these are good uses. Two students noted that they had not used it much but think they would like to use it more.

 

I’m happy that the majority of the students are using the study guides at least semi-regularly and seem to be benefiting from them. I was expecting the varied use, and I don’t see that as a problem. In my drop-in (office) hours, I’ve been able to encourage the students to use the guides more regularly, and several students regularly come by to clarify their answers to the study guide even if there isn’t an exam coming up. That’s been a good thing. I purposefully didn’t give answers (other than numerical ones) to the questions to encourage students to come in and talk through their answers with me.

 

Q2: What in the study guide have you found helpful?

 

By far the “test yourself” questions were what students found most helpful. They also used it as a summary of the key points of each class. The students commented that they liked being able to see more examples of how I ask questions, in contrast to the online homework that is more limited in terms of conceptual pieces. A couple of students mentioned liking the “real-world applications” that accompanied some of the questions. I admit that I’m haphazard. Sometimes I come up with an interesting question, and sometimes I don’t. Depends on how much time I have to work on it. I expect to improve the guides over time. One student specifically mentioned that my format of asking questions (rather than making a statement) to reinforce important conceptual points was helpful. Many of them also used the study guide to organize the key takeaways from each class meeting. By and large, the range of responses covered what I expected.

 

Q3: What in the study guide could be improved?

 

Students wanted more detailed answers and worked solutions step-by-step. I was expecting this response even though I gave the rationale for why I wasn’t providing detailed answers. Students wanted to check if they were on the “right track”. I wanted them to do that check with me, rather than reading (and potentially parroting without understanding) a written answer I provide to a conceptual question. Also, because there is often more than one way to answer a question, I don’t want students to get overly narrow in the way they tackle problems or answer questions. At the moment, I’m sticking with my plan not to provide detailed answers and see how things go, although I might be convinced otherwise. The main reason is that the current approach encourages students to drop-in and talk chemistry, and I think this really helps them learn better more so than reading off whatever answers I might provide. (I do worked examples step-by-step in class, and the textbook also does this.)

 

There were a couple of comments about my questions (be they conceptual or “test yourself”) being vaguely worded. A student says: “… makes it hard if you are trying to pin-point what you don’t understand”. This makes sense. If you don’t know the material very well, which will be true for a learner, then I could see how I need to be more careful or clearer when I phrase my conceptual questions or points. Just because it’s clear in my head when I write a question doesn’t mean that it won’t be misinterpreted by the student who thinks I’m asking for something else. (I also see this on exams where a student writes an answer to a question I haven’t asked, although much less so for this strong Honors class.) Anyway, this tells me that I need to spend a bit more time thinking about how I’m phrasing things in the study guide. I hope that several more iterations will improve this.

 

Q4: What could I do as the instructor to help students use the study guide more effectively?

 

Several students mentioned that more reminders to use the study guides would be helpful! One mentioned having a short guide of “how best to use the study guides”. I talked about this in class, but have little verbiage on my course website – I should include more verbiage and that might help. One suggested that “students may be more incentivized to use the study guide if they knew that the quiz questions were being pulled from the study guide or if test questions were likely to mirror those that are on the study guides”. That’s a good idea. I had mentioned the second, but not the first. I could emphasize this point more.

 

One student mentioned that the study guides could be intimidating because it looks like there’s so much material. Very true. Turns out there is a lot of material in chemistry to learn. I’m using the current “daily” format to break it into small chunks. In a previous iteration, the study guides were more voluminous and even more intimidating. (The present students don’t know what the previous iterations looked like; one student mentioned being curious about “how the setup is different or changed from previous semesters”.)

 

Several students thought nothing needed to be changed. So I’m doing at least some things right. Here’s one student’s response to the question: “Nothing! They are extremely useful in that they emphasize the most important elements of each lecture session. I believe that [the professor] has provide all of the tools for success and it is up to students to determine how they will utilize these tools throughout the course.”

 

Overall, I think the study guides are doing what I’d like them to do. The issue of whether to provide more “answers” will continue to be in tension with my pedagogical goals and what I believe best serves student learning even if students think a different approach will work better. Sometimes I take what they say to heart and make a change. Sometimes I don’t make a change. Regardless, it’s good to get the feedback.