Thursday, May 23, 2024

Walking 30 miles from Beverly to Boston

On May 21, 2024, I walked about 30 miles from Beverly MA to Boston MA.

A coworker asked me to blog about it, so here we go.

Here's a Google Photos album of pictures I took along the way:

https://photos.app.goo.gl/NkyMJwm3MSt64maD6

Here's an interactive map of the route:

https://www.google.com/maps/d/edit?mid=1yGUECeeST6aukYuqzeBJFqhrT6GI1tA&usp=sharing

And the record from Strava:



Saturday, December 31, 2022

More correspondence about religion

After 6 years of suspense (see my first comment on this post: http://breathmintsforpenguins.blogspot.com/2016/06/musings-on-meaning-of-meaning-and.html), my thoughtful Christian friend got back to me and we continued for a while brought it to a thrilling conclusion.

I thought it was a fun conversation, so I'm posting (a somewhat edited for brevity and clarity) my part of it here. I haven't asked for permission to post my friend's comments, so I'll summarize those. I guess if they come across this and want their part included verbatim, I'll replace with the full, unedited correspondence. Also skipping a big chunk towards the end because it gets complicated, and probably boring to read for people other than the participants.

Monday, January 4, 2021

Backpropagation through graphical models for new year's resolutions and planning out priorities for the next year

Around New Year's every year, I take some time to reflect on the successes and failures of the past year and decide what I should prioritize for the next year. Usually I do this using a combination of paragraphs and bullet point lists.

I recently took a deep learning class where I learned about a technique called backpropagation, which is used to find optimal parameters for a predictive model. It occurred to me that an analogous thought process could be used to help guide New Year's planning.

This is the first time I have gone through this exercise, so I'm making it up as I go along. If you decide to try this (or if you have already tried something similar), I would be really delighted to hear any feedback (good or bad) or suggestions you might have for improving the method. 

Friday, January 1, 2021

Dreaming of Grandpa


memorial for grandma

Productivity through cycles of exploration and focus


I'm not a very good swimmer. I don't usually wear goggles, but I don't like getting water in my eyes. When I swim I start by looking around, noticing where obstacles points of interest are, and picking a destination. Then I close my eyes and swim towards that target for a while before stopping and repeating the cycle.

Isolating (peppermint) glandular trichomes by the bead-beater method

Sean Johnson

April 26, 2017

Simplified from: Gershenzon et al., ANALYTICAL BIOCHEMISTRY 200,130-138 (1992)


Tuesday, July 9, 2019

Zamenhof 1907 speech at the London guildhall

Zamenhof gave a lot of beautiful speeches. This one makes me cry every time I read it. It's a lot more eloquent in the original Esperanto, but here is a translation into English that I've made.

The speech was made during a period of intense nationalist violence in Zamenhof's home region. Also at the time of the speech, some members of the Esperanto community thought that Zamenhof's idealism was scaring people away from Esperanto. They wanted to advertise and use Esperanto solely as an instrument for international business and communication and to drop any discussion of world peace. A few years earlier Zamenhof had reluctantly endorsed some of their ideas.

In this speech Zamenhof clarifies that he still sees the pursuit of world peace as the primary goal of Esperanto and he makes a plea for people in the Esperanto community (and beyond) to not lose their idealism. For more about Zamenhof and his political ideas, see here, and here.


Saturday, June 8, 2019

If the AI alignment problem is impossible, then unfriendly superintelligent AI might be still be self limiting

I thought of the following idea while trying to rationalize in-universe explanations for why evil AI in science fiction stories (specifically The Matrix) haven't reached the point of singularity and can still sometimes be outsmarted by humans.

It is possible that unfriendly AI would be self-limiting in the sense that it would voluntarily choose to not create an even more powerful unfriendly AI. For example, if an unfriendly AI were just smart enough to usurp humans as the dominant intelligence on Earth (and either destroy or subjugate humanity), then maybe that AI would want to avoid making the same mistake that people made and would retain its dominant status by choosing not to create something more intelligent than itself. It could be that the AI alignment problem is so hard that a superintelligence can't be aligned, but a superintelligence may be smart enough to realize that it (and more advanced superintelligences) can't be aligned.

Friday, January 25, 2019

to postdoc or not to postdoc: from PhD to industry

After finishing my PhD, I did an academic postdoc for about 2 years, then recently moved to a small company in private industry. A friend recently asked whether I thought the academic postdoc was worthwhile and if I had any advice for getting into industry. Here's what I said.

Saturday, January 12, 2019

How systems programming is like biochemistry


Recently I've been taking a lot of classes in computer architecture and systems programming, someone asked me why, with a biology background, I would be interested in such things. Here's what I said.

Thursday, January 3, 2019

compiling Codon Optimizer on ubuntu 18.04


There is a cool old piece of software for codon optimization, inventively named "Codon Optimizer"

http://www.cs.ubc.ca/labs/beta/Projects/codon-optimizer/

It's not really very useful from a practical standpoint because its codon optimization strategy is to try to replace every codon with the highest CAI codon for the same amino acid. Which turns out to not be a very good strategy in practice. The software also can only optimize for E. coli, and has the translation tables and codon frequency tables hard-coded in.

Nevertheless, it's one of the few Open Source programs for codon optimization, so it might be useful to use as a starting point for a more fully-featured program, by adding more tables or some alternative algorithms for picking codons to use.

Unfortunately, it won't compile on Ubuntu 18.04.

To get it to compile, I had to do the following:

Minimalist shoes: a journey


The idea of minimalist footwear appeals to me from a number of perspectives. Aesthetically, I like the idea of having as little as possible on my feet. Minimalist shoes also tend to be cheaper than other shoes, particularly other running shoes. I also like the fact that I can (try to) make my own minimalist shoes, whereas trying to make a pair of padded, heel-raised shoes seems like it would be a nightmare.

Thursday, December 13, 2018

Codon translation tables in JSON format

I reformatted NCBIs genetic code tables (https://www.ncbi.nlm.nih.gov/Taxonomy/Utils/wprintgc.cgi?chapter=cgencodes) into JSON format.

Check out this link (http://www.petercollingridge.co.uk/tutorials/bioinformatics/codon-table/) for a quick algorithm for parsing these strings.

 codon_tables = {  
 "1": ["FFLLSSSSYY**CC*WLLLLPPPPHHQQRRRRIIIMTTTTNNKKSSRRVVVVAAAADDEEGGGG","---M------**--*----M---------------M----------------------------","Standard"],  
 "2": ["FFLLSSSSYY**CCWWLLLLPPPPHHQQRRRRIIMMTTTTNNKKSS**VVVVAAAADDEEGGGG","----------**--------------------MMMM----------**---M------------","Vertebrate Mitochondrial"],  
 "3": ["FFLLSSSSYY**CCWWTTTTPPPPHHQQRRRRIIMMTTTTNNKKSSRRVVVVAAAADDEEGGGG","----------**----------------------MM----------------------------","Yeast Mitochondrial"],  
 "4": ["FFLLSSSSYY**CCWWLLLLPPPPHHQQRRRRIIIMTTTTNNKKSSRRVVVVAAAADDEEGGGG","--MM------**-------M------------MMMM---------------M------------","Mold Mitochondrial; Protozoan Mitochondrial; Coelenterate Mitochondrial; Mycoplasma; Spiroplasma"],  
 "5": ["FFLLSSSSYY**CCWWLLLLPPPPHHQQRRRRIIMMTTTTNNKKSSSSVVVVAAAADDEEGGGG","---M------**--------------------MMMM---------------M------------","Invertebrate Mitochondrial"],  
 "6": ["FFLLSSSSYYQQCC*WLLLLPPPPHHQQRRRRIIIMTTTTNNKKSSRRVVVVAAAADDEEGGGG","--------------*--------------------M----------------------------","Ciliate, Dasycladacean and Hexamita Nuclear"],  
 "9": ["FFLLSSSSYY**CCWWLLLLPPPPHHQQRRRRIIIMTTTTNNNKSSSSVVVVAAAADDEEGGGG","----------**-----------------------M---------------M------------","Echinoderm and Flatworm Mitochondrial"],  
 "10":["FFLLSSSSYY**CCCWLLLLPPPPHHQQRRRRIIIMTTTTNNKKSSRRVVVVAAAADDEEGGGG","----------**-----------------------M----------------------------","Euplotid Nuclear"],  
 "11":["FFLLSSSSYY**CC*WLLLLPPPPHHQQRRRRIIIMTTTTNNKKSSRRVVVVAAAADDEEGGGG","---M------**--*----M------------MMMM---------------M------------","Bacterial, Archaeal and Plant Plastid"],  
 "12":["FFLLSSSSYY**CC*WLLLSPPPPHHQQRRRRIIIMTTTTNNKKSSRRVVVVAAAADDEEGGGG","----------**--*----M---------------M----------------------------","Alternative Yeast Nuclear"],  
 "13":["FFLLSSSSYY**CCWWLLLLPPPPHHQQRRRRIIMMTTTTNNKKSSGGVVVVAAAADDEEGGGG","---M------**----------------------MM---------------M------------","Ascidian Mitochondrial"],  
 "14":["FFLLSSSSYYY*CCWWLLLLPPPPHHQQRRRRIIIMTTTTNNNKSSSSVVVVAAAADDEEGGGG","-----------*-----------------------M----------------------------","Alternative Flatworm Mitochondrial"],  
 "16":["FFLLSSSSYY*LCC*WLLLLPPPPHHQQRRRRIIIMTTTTNNKKSSRRVVVVAAAADDEEGGGG","----------*---*--------------------M----------------------------","Chlorophycean Mitochondrial"],  
 "21":["FFLLSSSSYY**CCWWLLLLPPPPHHQQRRRRIIMMTTTTNNNKSSSSVVVVAAAADDEEGGGG","----------**-----------------------M---------------M------------","Trematode Mitochondrial"],  
 "22":["FFLLSS*SYY*LCC*WLLLLPPPPHHQQRRRRIIIMTTTTNNKKSSRRVVVVAAAADDEEGGGG","------*---*---*--------------------M----------------------------","Scenedesmus obliquus Mitochondrial"],  
 "23":["FF*LSSSSYY**CC*WLLLLPPPPHHQQRRRRIIIMTTTTNNKKSSRRVVVVAAAADDEEGGGG","--*-------**--*-----------------M--M---------------M------------","Thraustochytrium Mitochondrial"],  
 "24":["FFLLSSSSYY**CCWWLLLLPPPPHHQQRRRRIIIMTTTTNNKKSSSKVVVVAAAADDEEGGGG","---M------**-------M---------------M---------------M------------","Pterobranchia Mitochondrial"],  
 "25":["FFLLSSSSYY**CCGWLLLLPPPPHHQQRRRRIIIMTTTTNNKKSSRRVVVVAAAADDEEGGGG","---M------**-----------------------M---------------M------------","Candidate Division SR1 and Gracilibacteria"],  
 "26":["FFLLSSSSYY**CC*WLLLAPPPPHHQQRRRRIIIMTTTTNNKKSSRRVVVVAAAADDEEGGGG","----------**--*----M---------------M----------------------------","Pachysolen tannophilus Nuclear"],  
 "27":["FFLLSSSSYYQQCCWWLLLLPPPPHHQQRRRRIIIMTTTTNNKKSSRRVVVVAAAADDEEGGGG","--------------*--------------------M----------------------------","Karyorelict Nuclear"],  
 "28":["FFLLSSSSYYQQCCWWLLLLPPPPHHQQRRRRIIIMTTTTNNKKSSRRVVVVAAAADDEEGGGG","----------**--*--------------------M----------------------------","Condylostoma Nuclear"],  
 "29":["FFLLSSSSYYYYCC*WLLLLPPPPHHQQRRRRIIIMTTTTNNKKSSRRVVVVAAAADDEEGGGG","--------------*--------------------M----------------------------","Mesodinium Nuclear"],  
 "30":["FFLLSSSSYYEECC*WLLLLPPPPHHQQRRRRIIIMTTTTNNKKSSRRVVVVAAAADDEEGGGG","--------------*--------------------M----------------------------","Peritrich Nuclear"],  
 "31":["FFLLSSSSYYEECCWWLLLLPPPPHHQQRRRRIIIMTTTTNNKKSSRRVVVVAAAADDEEGGGG","----------**-----------------------M----------------------------","Blastocrithidia Nuclear"]  
 }  

Thursday, August 17, 2017

targetP wrapper for large queries

As far as I know, TargetP is still (17 years after its original publication!) the best software for predicting subcellular localization for plant proteins, and also the location of truncation sites.

Without any modifications, targetp works well with small (by modern standards) queries, of less than 2,000 sequences at a time. But becomes glitchy when running with larger queries, such as the 30k-100k genes that are typical from a plant transcriptome assembly.

To adapt TargetP for larger queries, I wrote a Python script that acts as a wrapper around TargetP, called targetp_all.py. The script works by separating the input into smaller subsets of sequences and running those, and combining the output.

Interface is the same as the original program but with a few additional options. The output is somewhat simplified to be in tab-separated format.

It would also be nice to be able to parallelize the execution of TargetP to run on multiple cores at once, but I haven't attempted this yet. I believe that there will be complications involving conflicting temporary files, that may require careful modification of the original source code.

Source code follows. BioPython is a dependency.

Saturday, April 29, 2017

Mira4 assembly of 454 reads from SRA

I want to make an assembly of the Annona squamosa fruit transcriptome data from this paper (http://dx.doi.org/10.1186/s12864-015-1248-3). They give in the paper a link to a web resource (http://www.annonatranscriptome.nabi.res.in/), but the resource appears to now be defunct, so to get contigs reads, I will have to assemble the reads myself. The reads are from two different cultivars of Annona squamosa, so I'm going to assemble each cultivar separately first, and then if that works, I'll try a combined assembly.

MIRA is a nice, free, software package that can assemble 454 data. I've had success with it before, so that's what I'll use for this project too.


Monday, March 20, 2017

Tips for Methods Development and Optimization in Biology



I haven't posted anything in quite a while. Mostly that's because I haven't written anything that I thought would be of general interest. I once again have a young lab assistant to preach at, so I might as well preach at the world too. Here is what I've written for her about methods development in the biology laboratory.

Thursday, November 10, 2016

The Loss of the Creature, Walker Percy, detailed commentary

I think the essay "The Loss of the Creature" by Walker Percy (from the book The Message in the Bottle) is well worth reading and thinking about. In this post I offer a detailed, paragraph-by-paragraph commentary of the essay. This post isn't really meant to be read from beginning to end (if you try that, it may get repetitive). It's meant more as a set of detailed footnotes for people who find the essay to be confusing. The way to follow this post is to print off Percy's essay, then number the paragraphs. By my numbering, there are 38 paragraphs in section I, and 24 paragraphs in section II (starting with #39, ending with #62). When reading, if you get stuck on a paragraph, look it up here, and maybe my comments will make it more clear (hopefully they won't make it even more confusing). For a more personalized discussion, please leave a comment. I'm well aware of the irony of writing an analysis of this essay as though I expect people to experience the essay through my interpretation (exactly opposite to how the essay encourages us to experience the world). I would encourage you to not think of my commentary as authoritative, but maybe just as a spark for your own thought.

(This is a work in progress, I got about half way through and then set it aside. It's been sitting unfinished for long enough that I feel I might as well just publish what I have. I hope to finish it eventually. In the mean time, if anybody else would like to contribute commentary for the remaining paragraphs, that would be nice.)

Wednesday, October 26, 2016

My dissertation is now available for download from ProQuest

My PhD dissertation is now available for free online at http://pqdtopen.proquest.com/pubnum/10164019.html

I'm also making my laboratory notebooks from grad school available. They don't cover everything I did, they're pretty messy, and I'm not sure they will be useful or interesting to anyone, but here they are.
book1
book2
book3


Here is the abstract to my dissertation:
Plant natural products are useful for many different applications, including medicines, flavors and fragrances, and industrial uses. Two important aspects of plant natural products research are the identification of compounds in their source plants, and the characterization of the processes involved in their biosynthesis. To aid in the identification of plant natural products, we developed the Spektraris family of databases. These databases include highperformance liquid chromatography mass spectrometry data, and 13C and 1H nuclear magnetic resonance data, which are searchable through an online interface. The utility of Spektraris was validated by using it to identify compounds in plant extracts and as part of a workflow to elucidate the structure of a previously undescribed compound.
Mints have a long history of use as model systems for studying the processes of terpene natural products biosynthesis in specialized plant tissues. The mint family (Lamiaceae), synthesizes and stores volatile terpenes in glandular trichomes. Using a comparative transcriptomic approach, we identified differences in gene expression of monoterpene biosynthetic genes among mint species with different oil profiles. We also assembled the genome of a mint species, Mentha longifolia. The genome assembly will be valuable for future mint research.
To further investigate biosynthetic processes in mint, I developed a detailed mathematical model of the metabolism of peppermint glandular trichomes. The model incorporates multiple sources of data, including transcriptome data, metabolite data, enzymatic data from the peppermint literature, and previously developed models of plant metabolism. The creation of a new metabolic modeling software package, called YASMEnv, facilitated construction of the model. Model-based simulated reaction knockouts using flux balance analysis revealed that fermentation may be important for ATP regeneration in secretory phase glandular trichomes. Follow up experiments confirmed high levels of alcohol dehydrogenase activity in secretory phase isolated trichomes. Simulations also supported an essential role for ferredoxin and ferredoxin-NADP reductase. Transcriptome analysis revealed the presence of an isoform of ferredoxin in trichomes distinct from the one expressed in root. The presence of a distinct ferredoxin isoform in trichomes supports the hypothesis that selection pressure for efficient natural products biosynthesis may also act on the enzymes of primary metabolism.

Thursday, October 13, 2016

rendering zdock server results using python and UCSF Chimera

Problem:
I was interested in learning where the protein-protein interaction sites were on a particular protein (the receptor). I had pdb files for that protein, and for three proteins that it is known to bind to (the ligands). There are specific amino-acids of interest on the receptor protein, where there is variability among different species. For every combination of receptor and ligand, I want to perform a prediction of where the proteins interact, and then generate an image where the variable amino-acids are highlighted.