Category: thoughts

  • Replace “using AI” with “using interpolation”

    (2nd try after HN feedback)

    When cloud computing was a new, cool buzzword thrown around that no one understood, a shortcut was suggested whether it is the right solution:

    In any sentence, replace “in the cloud” with “on another person’s computer”.

    “Computing on the cloud” becomes “computing on another person’s computer”.
    “I backup my data to the cloud” becomes “backup to another person’s computer”.
    This makes it easier to weigh benefits – they provide more space, and provide maintenance, against drawbacks – they have your data and fully control the computer it sits on.

    Today, “AI” is thrown around as an all-encompassing buzzword. “Enhance your start-up with AI”, “AI helps prepare lawsuit”, “AI-written lawsuit contains mistakes”, “Fighting wildfires with the help of AI”, “Get data insights with AI”.

    A shortcut to clarify thinking about AI

    To expose overuse of the term “using AI” without meaning, I propose to replace it with “using interpolation”: “Enhance your start-up with interpolation”, “Using interpolation helps prepare lawsuit”, “Interpolated lawsuit contains mistakes”, “Fighting wildfires with the help of interpolation”, “Get data insights with interpolation”.

    This makes a few things clear: First, there are the benefits of any automatic data processing approach; this is not new with AI. Secondly, the replacement is effective at stripping away the mythical meaning of “AI” as an independent actor, removing the possibility to delegate ownership to it. Saying “my interpolation did it” sounds today much sillier than “my AI did it”.

    Many, but not all AI methods can indeed be framed as interpolation, with a very complicated, high-dimensional distance function among outputs, to determine what to produce given an input. This distance function, or slatent space, was shaped from data.

    The interpolation framing reveals the first question to ask:

    1. Where do the underlying data come from, who made them?

    This first question leads you to problematic biases in the training data that the AI inherits. It can also reveal copyright issues and whether the data producers are fairly compensated.

    Now you are also more plainly seeing someone saying “I’m doing X with technique Y, ” which leads to the question: Is this better than what was there before? This is actually a two-part question:

    2. By what performance metric do the claimants want to be judged?

    This second question reveals what people value, and whether this aligns with your values.

    3. Is the performance better in that metric than the current method?

    Question three reveals whether there was an improvement made so far, and may help distinguish vaporware from genuine improvement. The baseline may be another AI method, for example comparing LLMs to Markov Chains trivially reveals how much better LLMs are. Only very few analyses truly lack a baseline.

    Conclusion

    Demand from AI articles to know the data origin, performance metric, and historic baseline.

    If they don’t give that information, replace “using AI” with “using interpolation.”


    Postscript:

    Note that I did not include “what is the model?”, i.e., the architecture or training procedure – this is the least important technical detail.

    Questions 1) and 3) are often not made by the claimants, but reused. This is often not or extremely briefly described, reflecting how much time the AI developers spent on it. These three questions are essential for putting the AI claim into context.

    Some cop-outs:

    • If only similar performance is achieved, the AI developers often point to improvements in processing time, which should be added to the performance metric answer of question 2. It’s a valid goal to achieve somewhat comparable performance at much faster speed.
    • My pet peeve cop-out is “it has potential for improvement in the future.” This may be technically true, because models might learn with more training data. However, at least in a scientific and policy context a opportunity for critically and objectively examining the outcome, in context of the historical baseline, is being skipped.
  • Perfectionism and Dilettantism

    Finishing projects is an art. You need to balance pragmatism of getting things done with the bar on quality you set for yourself.

    Many people I know are perfectionists, and have a high standard for themselves. This often comes from parents saying something like “why did you not get the highest grade?”, and class mates laughing when we do something poorly. Consequently, we are embarrassed to not deliver perfection. I guess parents ingrain this into children to embed a drive to their kids life, and to do their work well.

    A harmful consequence of this is that it is really difficult to pick up new hobbies and independent skills. Whenever you try something new, first you suck at it. It’s frustrating because even as a beginner, you can tell very much how bad you are at it.

    When I told my cousin, who installed a pull-up bar at home, that cannot do many pull-ups, he said to me: “That’s great! That’s the time when you can improve quickly!” But it is not easy to get frustrated, knowing that there are many people much, much better at what you are trying to learn.

    Whenever someone says to me, “I want to learn X? How can you be so good at it?”, I respond:

    How do you get good at anything? You do it shitty a thousand times first. — Johannes Buchner

    Being bad at something but still practicing it for the fun of it, without ambition to become the best, is called dilettantism. Be a dilettante in many things and ignore the perfectionism.

    I wrote this text, for example.

  • IQ is for stupid people

    In 2017, I took an Uber in Boston to the airport. The driver was a young Russian. We talked about our backgrounds, and he mentioned “oh, you must be very smart to be a scientist”. Then he brought up that there might be a few extremely smart people, in the world, like with an IQ of 400, which control everything because they see through it all.

    Intelligence is difficult to define, but one definition is the adaptability to solve new situations. My preferred definition is different, and comes from reading a book about sheep, who try to achieve a task but are hindered by getting distracted and forgetting. I’d say intelligence is the ability to hold and rationally follow a chain of thought in your mind for a sustained time.

    Intelligence can be trained and improved, but you also need the luxury of being free of distractions to focus and plan.

    Reasoning alone is not worth much though, you need knowledge. One example is the debate of Sam Harris and Noam Chomsky. Sam Harris tries to reason from first principles and definitions. Noam Chomsky, additionally, builds on deep background information about world history, political activities, and cultural contexts. Sam Harris tries to put this messy reality away to argue cleanly, and this fails to be convincing.

    So let’s imagine an individual with extreme intelligence and knowledge. Could that individual bend the world to their will? I think there are obstacles: People just do random things and are not consistently rational to be influenced. Influencing people’s behaviour is not easy. You need to be a social person (or a psychopath I suppose), some people call this emotional intelligence or EQ. But even then, people are defaulting on their habits and culture. Finally, this influencing does not scale. To change one countries’ policy, you need to build alliances spanning hundreds of people. Unless you are lucky and powerful people already want to do what you want them to do (but then what is your influence?), or you have hard work ahead of you.

    When I was a teenager, the Mensa organisation (IQ>130) was spoken about in a similar tone as the Illuminati (they were Bavarian by the way). I met some of them, they were highly into chess, quirky, and a bit on the spectrum.

    My main problem with IQ is that if you start talking about IQ, you have already lost: you are reducing your self to a single number and implicitly accept it as a way to quantify your worth. So I never took an IQ test. It’s dumb. I’m glad the world moved on from it.

  • Are rich people kinder?

    In the movie Parasite, the poor family going through a tough time is sitting together talking about the rich family they just interacted with:

    Ki-taek: She’s so naive and nice. She’s rich but she’s still nice.
    Chung-sook: Not “Rich but still nice.” Nice because she’s rich, you know? Hell, if I had all this money, I’d be nice too!

    The stress from having to perform all the time to bring food to the table can put people in a bad state of mind, where they are less able to take a breather. Dealing with difficult people then does not come easy. Shelling out bullshit money to sort something is not an option.

    I don’t think though that rich people are kinder overall, in terms of giving. There are so many people who spend their lifetime building up their communities and helping others, who will not be on any list of rich people. Indeed, the richest donate a small fraction of their money … it is rare for rich people to disappear from the list of rich people.

    The analogy has been made that billionaires are like Smaug the dragon, hording their money for doing … nothing useful. They are withholding money from the economy, money that does not circulate as well as if it was with poorer people. Money only trickles up and then gets stuck there, and only through taxes on income and/or wealth, this effect is being made less severe.

    It’s not a coincidence that equality leads to higher happiness, not just for poor people, but also for rich people who have then less, because they do not have to fear the strong economic difference and stronger cohesion.

  • Replace “using AI” with “using computers”

    When cloud computing was a new, cool buzzword thrown around that no one understood, a shortcut was suggested whether it is the right solution:

    In any sentence, replace “in the cloud” with “on another person’s computer”.

    “Computing on the cloud” becomes “computing on another person’s computer”.
    “I backup my data to the cloud” becomes “backup to another person’s computer”.
    This makes it easier to weigh benefits – they provide more space, and provide maintenance, against drawbacks – they have your data and fully control the computer it sits on.

    Today, “AI” is thrown around as an all-encompassing buzzword. “Enhance your start-up with AI”, “AI helps prepare lawsuit”, “AI-written lawsuit contains mistakes”, “Fighting wildfires with the help of AI”, “Get data insights with AI”.

    Is there a similar shortcut that clarifies thinking?

    To expose overuse of the term “using AI” without meaning, I propose to replace it with “using computers”: “Enhance your start-up with computers”, “Using computers helps prepare lawsuit”, “Computer-written lawsuit contains mistakes”, “Fighting wildfires with the help of computers”, “Get data insights with computers”.

    This makes a few things clear: First, there are benefits to using computers because of their automatic data processing. This is not new with AI. Secondly, the replacement is effective at stripping away the mythical meaning of “AI” as an independent actor, removing the possibility to delegate ownership to it. Saying “my computer did it” sounds today much sillier than “my AI did it”.

    So, replace “using AI” with “using computers”. It reveals to you how little the statement by itself tells you.

    To seriously talk about AI, we have to unpack the term. In general terms, we are talking about methods that expand their capabilities with increasing data. Therefore to judge whether AI is good or not for an application, we need to find out:

    1. Where do the data come from, who made them?
      • Question 1 leads you to problematic biases in the training data that the AI inherits. It can also reveal copyright issues and whether the data producers are fairly compensated.
    2. By what performance metric do the claimants want to be judged?
      • Question 2 reveals what people value, and whether this aligns with your values.
    3. Is the performance better in that metric than the current method?
      • Question 3 reveals whether there was an improvement made so far, and may help distinguish vaporware from genuine improvement. The baseline may be another AI method, for example comparing LLMs to Markov Chains trivially reveals how much better LLMs are. Only very few analyses truly lack a baseline.

    Note that I did not include “what is the model?” – this is the least important technical detail.

    Questions 1) and 3) are often not made by the claimants, but reused. This is often not or extremely briefly described, reflecting how much time the AI developers spent on it. These three questions are essential for putting the AI claim into context.

    Some cop-outs:

    • If only similar performance is achieved, the AI developers often point to improvements in processing time, which should be added to the performance metric answer of question 2. It’s a valid goal to achieve somewhat comparable performance at much faster speed.
    • My pet peeve cop-out is “it has potential for improvement in the future.” This may be technically true, because models might learn with more training data. However, at least in a scientific and policy context a opportunity for critically and objectively examining the outcome, in context of the historical baseline, is being skipped.

    Demand from AI news articles to know the data origin, performance metric, and historic baseline.

    If they don’t give that information, replace “using AI” with “using computers.”

  • Reliability

    Three decades ago, the word was “four nines”.

    Internet providers and company server administrators would take pride to deliver a service that was working 99.99% (four nines) or 99.999% of the time,. In German and Swiss culture, reliability is highly valued and this has been a differentiating factor, for example for automobiles and watches. Chains such as Starbucks largely build on your feeling of security of knowing exactly what you will get if you enter.

    Removing uncertainty from an uncertain life has value.

    A bit after the 2000s, reliability stopped being a focus for software services. Github, a key infrastructure component for many people to the point that people stop working when Github is down, does not promise four nines or make any reliability commitments. From EU, US to Chinese markets, towards the latter it is much more acceptable to customers that exciting new software is comes with bugs that need to be ironed out later. This results in quicker delivery timelines and innovation.

    So it is not necessarily a bad thing that it has become more acceptable that software and software services are unreliable. Building reliability takes continuous optimization over time and focused effort. The trade-off of reliability is dynamic adaptation to new markets and changed situations.

    Say for example a colleague sent you a file and its not quite the format that you can use. You write that short script or command line command (sed, awk) once, use it, verify that the output works, and never use that script again. You will not get exactly the same file format again. You do not need to support all possible variations of file formats that might arrive in the future.

    Data pipelines in research are much like this. It took me time to stop building generic multi-purpose frameworks. Most software developed in research projects is ad-hoc data transformations. By definition, research deals with something new. For the specific project, you mangle data from one form to another. You verify that it is all correct. The data pipeline specific to the project will likely not be reused in exactly the same way for the next project, because the next project is different, the exact form of the input or output, or the validity considerations will not be the same.

    Now elements of the script may be reused. We identify the commonalities. If many scientific projects benefit, the community rarely but surely invests into building a reliable, reusable data reduction pipeline. The process to get there is non-transparent and messy, but some truly impressive pipelines exist that qualify for four nines.

    For machine learning-based projects, reliability is both in and out of fashion. In principle, the starting point is great: Projects start by defining a metric to optimize (loss function). That is already a great exercise for defining what you care about, which is often otherwise left unspoken. That said, many AI-based research presentations leave me with the question: “Does it work?” and “Can I rely on it?”. Often the answer is no, formulated by the authors positively as “not yet”, and more specifically, with the phrase: “While it is not yet as good as previous approaches, our machine learning approach has the potential to outperform these in the future.” There is value in trying new approaches and it requires effort, often not achievable in a single 3-9 month project cycle. Maybe this is due to my cultural background, but I think we can be more ambitious on the standard we set for ourselves.

    You can see the lack of reliability also in a very basic metric: uptake. How many people use an AI-based tool productively? If it is not reliable enough, this will show up as a lack of citations.

    People might still talk about software that is not useful. On the one hand, this may be AI hype specific to a particular technique. This has also existed before. On the other hand, it may be genuine interest in a project in the future may opens a new type of analysis that has been accessible before. This is how new technology has always evolved – rather than improving established ways of achieving goals, focus on how new technology enables achieving different and new goals that could not be considered before.

    Despite being less visible, four nines are still important for non-customer facing backend software, but also for customer-facing software. If you want to dictate your email to your computer, or command a device, what error rate would be acceptable to you? If even 1-2 out of 100 sentences is wrong, and you have to work around the errors, the illusion of a smooth human-computer interface is broken. Take-up on voice-control is poor because it is more frustrating than a 100% reliability keyboard.

    For chat-based LLMs and AI agents, the threshold we accept seem to be astonishing low. You might get a useful answer only 20% of the time, but when it works the rush of excitement is enough to keep you going. The randomization is addictive, feeling like being at a slot machine for programmers. That said, makers of LLM optimize for exactly these performance reliability metrics, so we are seeing continuous improvements. The data driven approach means the reliability potentially reachable is certainly limited by the amount and purity of training data. The ceiling for reliability is, essentially, unknown.

  • The meaning of life

    Creating something is a form of expression. You put a piece of yourself into it. Effort, time, but also a piece of your craftmanship and thinking about how something should ideally be. For many artists, the act of creating their art is so important that giving it up would give up an essential part of their being.

    These days, software engineers are rediscovering just how important crafting software is to their identity. It has been true before that for many professional software engineers, their weekend hobby project was … creating software. For fun. It gives purpose to them to create something useful. Many times, the software is useful to exactly one person, and that’s okay.

    The advent of large text interpolators has created a scare among those that deal with text-based work. It is now quite possible to recreate software “from scratch”, if one has a technical specification, and a test harness. There are a few misunderstandings, however.

    1): Firstly, such a reproduction is a derivative work, because the LLM has been trained on all of github, often including the specific code and thus interpolating copyrighted work rather than extrapolating.

    2) Secondly, reproducing someone else’s work with LLMs does not build the skill to adapt the work to new specifications. Yes, you can adjust the prompt, but as anyone who works with LLMs knows while it feels cool and fast, it is neither a stable nor efficient workflow, you have to edit code yourself as well. Earlier this year I “lost control” over a software repository by letting an LLM make too many edits, to the point where I was unfamilar with the code and could not debug the mess any longer. I had to backtrack, delete large parts, and with discipline, built by hand the complex code that is not interpolatable from LLM training data.

    3) Thirdly, I have serious doubts many AI-created software packages will still be maintained after a year. Interest of the “creators” is largely not long-term in the domain, but rather with the AI toys.

    I still have hope that there ultimately is a feedback loop of “Has this person delivered value?” Not, “Has this person appeared to have delivered value?”, not “Has this person done something trendy?” but looking back at the created change in the bottom line.

    For me, there are three lessons:

    1) Creation gives purpose. From AI prompting pride in the creation cannot be generated, even if you give yourself titles like “Prompt engineer”. Effort is a key differentiating factor.

    2) Creation builds domain expertise. You can confidently argue about fine points. If you build with LLMs, your knowledge is shallow. Critical reading about a topic, including the background not directly relevant to a task at hand, is key here for retaining a big picture view of where a project should go.

    3) Computing and Society: Only by working and listening to people can you build something that matches a real need. One of the reasons why I departed from classical IT is that I felt many tools were built by software engineers for improving the lives of software engineers, but lacking a bigger goal. With AI, even if coding were to fall away as a skill, the skill to identify requirements through dialogs and embedding a solution with people remains an essential software engineering skill.

    I find research software engineering quite fulfilling. I build tools that are used by myself and my colleagues to investigate the Universe. I call them software telescopes. Could I find similar purpose in another domain? Probably. Could I find purpose in life without building something and putting my heart into it? Maybe not.

  • Garching

    Today, we walked 100 meters from home to the park, set up a inflatable couch, leaned back and watched the Perseid meteor shower 🌠 . It was beautiful. Garching is one of the very few places on Earth where you can see the night sky clearly, and work as an astronomer.

    Garching is a small city, but it is big enough that it has one of everything you need.

  • Co-authorship policy

    Co-authorship and acknowledgement are forms of recognition for direct contributions to a work. I want to elaborate on the policy I follow, influenced by the groups were I have worked.

    I essentially agree with this blog post, which is well-written: https://ramblingsofanecologa.wordpress.com/2020/11/10/who-should-be-a-co-author-rules-and-etiquette-of-academic-authorship/

    There are multiple aspects of how to contribute to a paper: “conceiving the idea, designing the study/experiment, collecting the data, analysing the data, writing the manuscript”.

    Acknowledging experiment design in astronomy

    It is sort of obvious that you should offer co-authorship to someone who gave you the idea for the study you are executing. Unfortunately, there are people in astrophysics who take data and ideas from others without credit. One second-hand horror story from decades ago is that a young postdoc was giving a presentation in CalTech about an idea to observe M-dwarf spectra for the first time, and laying out a strategy to achieve it. During question time, a senior researcher said: “That’s a great idea! I’m gonna do that.” Another second-hand horror story is that a member of the NICER instrument team used his privileged position to monitored ad-hoc transient observing proposals (where the data go public immediately) and scooped people’s work. This became so bad that in transient conferences, people would hide their source coordinates; yet the person once went up to a speaker saying “Don’t think that just because you hide the coordinates I can’t find your source.”. We must not let these behaviours proliferate. Do not hire these people.

    It is less obvious how to credit someone who enabled the data sets that your study builds upon.

    At the start of my career I have underestimated the effort people put in writing observing proposals. This experiment design work is crucial and enables scientific studies. So if the data is public, it is professional to ask the PI of the observing proposal, if they have not published a paper on the observations yet, whether they would like to be co-authors. This would probably not apply if it has been more than 2-3 years since the data were taken, but funding issues, or a PhD student dropping out can delay data exploitation significantly. If you know the PI, you should always ask.

    I do a lot of work on relatively old, archival observations, and that there is freely available, well-calibrated data is one of the best things of astronomy. I have acknowledged these effort with extensive citations. Here are two examples: (1) In a huge archival study of 900 gamma-ray bursts discovered by Swift, I cited over 165 Astronomer Telegrams, because they contain follow-up observations and key quantities like redshifts. (2) In a recent SED fitting paper analysing data from ultraviolet to infrared based on spectroscopically identified quasars, I spent 1.5 pages acknowledging the server infrastructure and all the surveys:

    Acknowledging project-specific contributions

    There are nice guidelines by some journals, that I agree with. Unfortunately, there does not seem to be a astronomy journal giving co-authorship guidelines, or at least I have not seen it.

    My favorite guideline write-up is by the medical journal ICMJE, but the recommendations from Springer Nature are also similar):

    Co-authorship requires:

    1. Substantial contributions to the conception or design of the work; or the acquisition, analysis, or interpretation of data for the work; AND
    2. Drafting the work or reviewing it critically for important intellectual content; AND
    3. Final approval of the version to be published; AND
    4. Agreement to be accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.

    The forth point means, that a co-author should be able to defend the paper in a conference discussion, when someone says the study is garbage. If they don’t want to have their name on it, don’t have your name on it.

    The third point implies you need to get the okay of the person to submit it with their name on it, and indeed you absolutely cannot submit a paper to a journal with people’s names without them agreeing to be co-authors. So if you are not careful, you can end up in a difficult limbo, where you have put a person’s name on the co-author list, shared drafts, they never reply, and now you cannot submit. People in permanent positions can keep you, without a permanent position, in this state forever. That’s why, in my opinion, you should never put people’s name into the co-authors list in drafts before they have sent you comments, and thereby implicitly or explicitly saying they would like to be coauthors. That is, I believe in a opt-in co-authorship policy, rather than a opt-out policy, because it avoids the stalemate. It creates an incentive to say yes and send comments.

    For the second point, drafting the paper, you can help me a lot. Pointing out relevant papers to cite is a crucial contributions – literature review is a lot of work if you have the ambition to be thorough. Going beyond criticising text and providing improved text blocks is a crucial contribution for someone like me who has a hard time writing. I’m gonna love having you as a co-author if you do these, even if you did not touch the data or were in any discussion meetings.

    Regarding the first point: Indirect contributions, such as holding workshops, making software and tutorials publicly available, are not part of acknowledgements or co-authorship. Maintaining the IT infrastructure of the institute is important to enable this and other project, but part of the affiliation as an implicit acknowledgement. Project-specific scientific software support that enables smooth project completion should be acknowledged. You can be as generous as you want in the acknowledgements! There is no cost and it only helps generate good will. Acknowledge reviewers, hallway discussions and encouragements. Custom code development enabling the project should count towards co-authorship.

    Managing expectations: be upfront about your policy

    You do not have to agree with my co-author ship policy, and I do not expect you to have my policy. But: be up-front about what your policy is. What threshold do you apply for co-authorship?

    This can avoid frustration and conflicts by being upfront. I had two collaborations where I spent time in many meetings discussing the paper data analysis and interpretation, providing references from a literature review, and thought that I would be a co-author but the first author did not. Both cases were first authors from (astro)particle physics – I learned that they have a much higher threshold for co-authorship.

    Personally, I think that is dumb: There is essentially no cost to including someone as a co-author. It is encouraging to continue collaborating. Can you imagine my excitement when I was included into the list of contributors to the VLC project? I became and remain a hard-core fan of the project. So err on the side of generosity. I realise, however, that telling people to be generous is futile; you can only grow the feeling of generosity by experiencing wealth, kindness and generosity from others setting an example. The counterargument that <=3 authors allows all to be named in the citation is, to be frank, pathetic.

    In my most recent paper, I copied the ICMJE guidelines into the latex/overleaf as comments just above the co-author list, to be upfront. I kept a list of co-author names and their substantial contributions, but commented out until they gave comments.

  • Understanding exponentials: COVID19 and climate change

    Exponential growth is unintuitive. The default heuristics of our brains is that things will continue as they are, or keep growing linearly.

    In the most intense phases of COVID19, exponential growth of case numbers, and thus hospitalizations, and the trigger-happy reaction needed to stop this exponential growth curve early was understood by most people.

    In the case of climate change, there is a similar phenomenon. We consider the warming curve of 0.x degrees per year, and it seems linear. However, the number of extreme heat days, and the number of extreme weather phenomena, are the tail end of a distribution that is widening. Their numbers is growing exponentially, and with it the cost, both monetary, of human lives, and life quality – the fraction of the year where it is safe to go outside into nature may shrink.

    It is easy to become blackpilled and say it is all pointless because there is too little political will and action. Exponential climate change damage is extremely bad. Have a look at the details and compare the +2° +3°, +4°, +5°, +6° scenarios in the IPCC report (2023, 2014).

    But there is a uplifting, positive side for exponentially bad impact.

    Small actions, including and especially in the later stages past +2° warming, have a huge benefit in avoiding an even worse outcome.

    Don’t give up.