Blog

  • A lesson for mastering sports

    I have learned one lesson in one sport, which I have found to be universally applicable to most other sports.

    The lesson is as follows:

    Imagine your shoulder points and your hips forming a rectangle. This rectangle needs to be mobile. Start any action by moving that rectangle. Not flailing the arms and legs. Move that rectangle and be agile.

    Being quick and active with your hips, and engaging with your core is also a lesson I have learned in jiujitsu. I have found it applicable in:

    • table tennis
    • volleyball
    • basketball
    • snowboarding
    • windsurfing
    • martial arts

    and bringing it up to others, they have confirmed it also for tennis.

    Which sport do you play and does it apply there?

  • Astronomy bucket list

    🗹 get a paper published
    🗹 get a paper rejected
    ☐ publish a letter
    🗹 publish a research note
    🗹 publish in A&A
    🗹 publish in ApJ
    🗹 publish in MNRAS
    🗹 publish in PASP
    🗹 publish a software package
    🗹 publish a catalogue to vizier
    🗹 play ArXiV astro-ph localtime roulette
    🗹 referee a paper
    🗹 discover a new observational relation
    🗹 publish analysis of cosmological simulations
    🗹 do a population study
    🗹 have a observing proposal accepted
    🗹 have a observing proposal rejected
    🗹 have a TOO proposal accepted
    🗹 have a DDT proposal accepted
    🗹 have a funding proposal accepted
    🗹 have a funding proposal rejected
    ☐ have a funding proposal for students or postdocs accepted
    🗹 supervise student BSc thesis
    🗹 supervise student MSc thesis
    ☐ supervise student PhD thesis
    🗹 do observations in optical
    🗹 do observations in radio
    🗹 do observations in X-rays
    🗹 work for NASA have NASA work for you
    🗹 single-sentence abstract
    🗹 give a conference talk
    🗹 give a colloquium




  • How to get emotionally attached to software but leave no impact to science

    How do people get emotionally attached to software?

    I mean positively emotionally attached to the point of plastering stickers on their laptop, or negatively attached to the level of raising their voices in discussions.

    Here I explore a few causes that I have seen.

    1. Attachment to strict teachers

    This cause of attachment is common with compilers (Rust, Stan) and program provers. They nitpick and force you into a lane. After significant investment and boxing your problem into the right shape, they are given the green light like a priest giving you a blessing. You are now a member.

    The harsher the teacher, the more rigorous the trial, the more attached is the feeling of belonging to an elite club of survivors.

    2. Attachment to what you have spent building

    Just spending a lot of time building something causes you to be attached to it. Firstly, design decisions can be personal expressions. Secondly, after spending many long nights or even months working on it, there is a sunken cost. It’s hard to consider that you have “wasted” your time. If someone tells you, “why don’t you use tool Y”, you’d laugh it off, not even trying to learn tool Y because you understand your imperfect work of love.

    A more open minded point of view is to think of the “sunken cost” as an investment to deeply understand the challenges in that domain. Even if the developer drops the tool and switches over to something else, they now know the concepts and challenges deeply, at an expert level.

    Unconvincing value propositions

    Building new software is fun! Because of the two points above, the author is likely to proselytize around, trying to convert people to their great tool.

    The expressed value propositions can harbor a communication mismatch: The inside view of the developer is not the outsider view. Their values of “it is well understood” and “easy to change” do not translate to the next person who has not edited the software yet. Likely, the outsider has also seen a few people proselytize, and then the difference in value may be difficult to see.

    Work without impact to science

    Every week there are papers on arXiv where someone invested months of work on a new data analysis technique and finally finds that it does not outperform the classic benchmark.

    Sometimes this is stated outright. I think it is great science to try something new. But there are other cases.

    Sometimes the benchmark is a setting unlike how the analysis will be applied in the future (for example: complete data, spectroscopic redshifts, known noise), or the benchmark statistic is non-standard to the field. In any case, whether the new technique is better or worse, the abstract invariable ends with “And therefore, our technique is a promising platform for future data sets.”

    It is disheartening to think of the many person months wasted, because people did not think longer about what they should optimize for.

    Personal reflection

    I think I also suffer from some sunken cost w.r.t. nested sampling. It is quite addictive for me to think about algorithm variations.

    That “no one uses my software” is not a problem I have though. What keeps me grounded is seeing a clear need in conference and hallway dialogs, combined with critical thinking to crystallize out what the core question is.

  • AGN physics open questions

    How are black holes created?

    What is their spin distribution?

    What is the shape and location of the hot X-ray corona? We know it is compact (<10Rg, from microlensing) and parallel to the disc (from polarimetry). But how is it powered? In practical terms: what sets the LUV-LX relation? The relation can be reproduced with AGNSED models, but this requires the corona to change with the accretion rate. What is causing this? Why is the dependence with luminosity and not Eddington ratio, which is different to all other dependencies?

    What is the shape and location of the “soft excess”? Is it the Comptonized disc? Why does the Comptonization radius depend on Eddington ratio but not luminosity?

    How much central engine diversity is there, or do fundamental parameters (MBH, lam_edd) set the geometry? Broad line variability studies suggest there is geometric diversity.

    What sets the LMIR-LX relation? Why do the X-rays become relatively X-ary weak at high luminosities? Why is there a tight relation to 6µm and 12µm instead of vast scatter? Is it because whenever the AGN is on, there is likely dust nearby?

    Are all bolometric corrections Eddington-ratio dependent? Does the LUV-LX relation bend over at high luminosities? The LMIR-LX does, so it would make sense if the LUV-LX does as well. How does the fundamental plane of black hole activity (LX-Lrad-MBH) play along here?

    What causes the covering factor to change with Eddington ratio, and why does this happen at different redshifts at different Eddington ratios? If the Eddington limit, modified to account for dust instead of just electron scattering, shapes the nuclear region (how exactly is it regulated? what is its shape? what is its distance?), then it is puzzling why this nuclear process should be limited differently at different redshifts. I suspect that instead, the regulation of the Eddington limit is due to the availability of gas, setting both covering factor and luminosity distribution. Like a small gas tank on a car with exponential hunger over time, the gas is unlikely to be full when power is at the maximum.

    To think a bit more systematically, here are the eleven components:

    1. Black hole – MBH, lam_edd
    2. Hot corona
    3. Warm corona
    4. Accretion disc
    5. BLR
    6. Hot dust
    7. Cold torus
    8. Outflows (UFO, atomic, ionised, molecular)
    9. Nuclear attenuation
    10. Host stellar potential
    11. Galaxy attenuation
    12. Host star formation

    1 is related to everything <8 by the Eddington ratio being important.

    4 is related to everything as a power source.

    are 2&3 related? lags perhaps suggest so.

    5 may be influenced by 2&3 through over-ionisation.

    6&7 is related to?

    4 -> 8 -> 10 -> 12 is a key process, potentially with involvement of 3 and 9.

    No funding agency cares about obscuration and the torus. Framing via fundamental black hole physics is essential.

    Transients probe accretion processes in a unique way, but the fueling is different and may not be the same because of the magnetic field structure. How to stand out in 30 years of AGN research?

    To test the empty fuel hypothesis, need SPHEREx + Euclid + eROSITA + SDSS-V.

    To test geometry of corona, need eROSITA + SDSS-V.

    To test co-evolution, need {eROSITA, NewAthena} + LS10 + Euclid + {SDSS-V,4MOST}.

  • Understanding AGN and their host galaxies

    Everyone and their dog makes a Spectral Energy Distribution (SED) fitting code. It’s not hard. The data are a few numbers, you just need Gaussian statistics, and connect to a model library.

    The design choices vary wildly though:

    Computation and uncertainty quantification

    cigale is a community platform mostly focused on galaxies, with quite limited uncertainty quantification abilities. Their grid computation method faces the curse of dimenstionality and consequently allows only a handful of free parameters. Because of this, it is essentially unsuitable for AGN. The better computational approach pursued by AGNFitter and FortesFit is to at least interpolate two grids (galaxy + AGN), and scale them linearly (M* and LAGN, respectively). This escapes the curse of dimensionality. GRAHSP computes SEDs without a grid, on-the-fly.

    The scalable approach to millions of objects is precomputing model photometry sampled from the parameter space and creating a neural network emulator. This is a lowd-to-lowd mapping, so quite feasible.

    Another approach is simulation-based inference.

    Galaxy model

    Galaxy archeology focuses on reconstructing the star-formation history (SFH). For this, a recent hubbub was the change from parametric SFHs to non-parametric variants, including with Gaussian processes (see for example Iyer et al, 2019). For AGN hosts, we are usually quite happy if we can figure out whether a galaxy is there (M* lower bound) and whether it is red, green or blue (which in any case, will be attributed to AGN feedback). If you read the papers advocating non-parametric approaches closely, they actually do say that for estimating the SFR in the last 100Myrs, it does not matter what you use.

    The more recent hubbub are binary populations (BPASS model), especially in young and low-metallicity environments seen in the UV. This is driven by high-z JWST spectra and photometry, but probably not relevant for z<6 AGN given current and future data quality.

    So most of the time probably something simple suffices, such as a delayed-exponential plus burst. This is quite at odds with what galaxy researchers want to do, and thus – multiple codes.

    AGN model

    If you then look at AGN physical models, there are only a few in existence and use:

    • for the accretion flow, Kobuta&Done’s AGNSED. seems to be most popular these days, which has the benefit of producing X-ray and optical emission.
    • for the torus: Fritz+06, CLUMPY/Nikutta08, SKIRTOR, CAT3D
    • there is essentially no good physics models for lines (broad, narrow, Fe forest).

    In my opinion, if you are interested to study component X of a system with N components, you better model the N-1 other components correctly, otherwise you will skew your inference about the component X. So you need empirical models for the N-1 components.

    Looking at empirical models, most common are Dale+14’s templates and, more recently, Temple+21 (broken powerlaw accretion disc, hot dust black body, broad and narrow lines, Fe forest, Balmer continuum), which has more parameters. GRAHSP’s model is similar in spirit, with a bending powerlaw accretion disc, hot and cold torus, broad and narrow lines, Fe forest, Balmer continuum). GRAHSP also models variability.

    If we believe we have an empircal model flexible enough to describe the entire AGN population to the quality of data available for the next few years, the next step is to bridge the gap between empirical models and physical models.

    The AGNSED model is a good starting point, because it models accretion physics. We just lack all the other parts:

    • Broad lines and Fe forest: we are far from a self-consistent model with enumerable parameters. This has to remain an empirical model.
    • Hot dust: this is essentially a black body.
    • Cool dust: probably enough to model with a few parameters.
    • Si feature: since this can be due to the host orientation, it is not worth modelling self-consistently.

    -> ML framing: Can we use SDSS spectra to derive a broad line and Fe forest model? Ideally, as a function of AGNSED parameters? For this, the classic starting point is to identify line-free portions of the spectrum, and fit AGNSED. We also know the MBH and FWHMs of the best lines (Ha, Hb, MgII), which we tie to the AGNSED parameters. This assumes single-epoch black hole mass estimates hold. Then we need an autoencoder to, conditional on the continuum level, predict the remaining mess of the emission spectrum, and the latent space holds the unknown physics. For the architecture, probably something like spender is necessary. It would be best to preprocess by subtracting the AGNSED continuum fit and remove attenuated and galaxy-dominated AGN.

    Biases in AGN morphology

    Galaxy morphology is very pretty to look at, but in terms of physics, mostly lacks useful information and is very difficult to analyse in a reproducible fashion.

    There are a few machine learning tools which use neural networks to classify morphology. Unfortunately, essentially all of them lack tests how the morphology indicators are impacted when you have a AGN in the center. So I do not trust them.

    Such tools can be useful for detecting weirdos for follow-up, but that’s not reliable demographics, and we need demographics to escape the curse of AGN flickering.

    -> ML framing: Can we carry over morphology information? Probably, if we use the contamination tests to quantify the impact of the morphology tool. The alternative, training a ML tool on images with contamination would be better, but likely we would still have false positives and false negatives (on morphology) to deal with.

    Unbiased Spatial+SED fitting

    There are a few approaches now which attempt resolved SED fitting. However, essentially all of them suffer from systematics because their morphology model is incomplete. You can see this by noticing that the decomposition is very fragile across bands, while flux errors are tiny.

    If you can reliably detect an extended galaxy component, you are in luck though, and able to break SED fitting degeneracies with that extended flux lower limit (EFLL) to the galaxy component. I think this will be powerful for Euclid.

    -> ML framing: Grant Stevens used the fact that AGN are the minority of galaxies to identify “AGN” as outliers in their central morphology. In this operational definition, any brighter than usual nucleus is an AGN. Similarly, we can use normal galaxies, and insert a blue-ish point source into the center and call that new hybrid object an AGN.

    If you go the empirical route, starting with an AGN population for the blue-ish point is a bad idea, because then separating out will be confused by the host galaxy the AGN carried with it. Instead, you can start with isolated stars and paste them into the center.

    For Euclid this should work, by selecting low extendedness objects and working in VIS, then training a ML tool to look at cutout images and predict the aperture flux without the central source. The loss should be punished for over-predicting the flux.

    Unbiased AGN host morphology

    Can we infer reliable morphology information? Probably, if we use the contamination tests to quantify the impact of the morphology tool. The alternative, training a ML tool on images with contamination would be better, but likely we would still have false positives and false negatives (detecting mergers for example) to deal with. The proper way to work here is to assign “Unknown” as a label when the AGN makes morphology unreliable and/or incorporate the skewed inference response into the analyses. No one is doing this properly though.

    One of the simplest properties of galaxies is their compactness, quantified as half-light radius. From the curve of growth, one could in principle read this off even for AGN hosts, but the half-light is not known if you do not know the total.

    Selection effects

    Lastly, quantifying the properties of members of a sample tells you about that sample and its selection, but not directly about the underlying demographics. There are many papers with Venn diagrams, but these hold little information because it remains unclear whether the missing subsets are expected or unexpected considering the depth of the selections.

    So it is essential to quantify selection effects. Ideally, you would analyse multiple selections and verify that you get the same results. A single generative model that creates populations of X-ray selected, optically selected, infrared selected, variability selected AGN is still missing. This is a necessary step to understand missing and disjoint populations.

  • The state of nested sampling research in 2026: II) Fundamental algorithm research

    see also part I.

    Nested sampling as an algorithm has a few papers presenting convergence proofs. It’s interesting from how many angles people have looked at this problem. See the relevant section in the literature review.

    Diagnostics (ongoing)

    Andrew Fowlie proposed an elegant test for diagnosing live whether nested sampling is sampling in an unbiased fashion, by tallying where in the likelihood-ordered list of live points the replacement point lands (I tend to confuse the wods order and rank). A variant is the statistically more powerful U-test ( section 4.5.2. of Buchner 2023). But ultimately, neither of these are as sensitive as you would wish.

    In practice, the most powerful technique is still observing the log(Z) value across reruns, e.g., when increasing the number of steps of the slice sampling procedure.

    More research is still needed on diagnostics. I suspect that the gradients of the dead points can be compared to the growth of the L-V curve and that could be more powerful.

    In the meantime, I transferred MCMC diagnostics to nested sampling, specifically the Jump Distance which traces the first-order autocorrelation length of a random walk. This research appears to be trivial for statistics journals and too non-astronomy for astronomy journals.

    Plauteaus (solved?)

    When the likelihood has plateaus, the original algorithm had an issue. In practice, this mostly occurs if your log-likelihood function returns a fixed number such as -1e300 when the parameters are invalid. Pragmatically, one can make that penalty slant towards the good region. More generally, an auxiliary with a tiny likelihood slant is a solution to plateaus, but in any case the nested sampling variant where all points with the same lowest likelihood value are removed has been proposed and should be the standard nested sampling method people implement (it is in UltraNest).

    Live point inheritance

    In most implementations, the dead point is replaced by sampling a new point, starting from another live point. This is true for ellipsoidal nested sampling and MCMC-based replacement schemes.

    In multi-modal settings, removing points can remove all points of a mode, and this was discussed early in the literature, mostly casually in conference proceedings (which I cannot find anymore, help!).

    Recently, I used analysis techniques from genetics to analyse the occurrence probability of mode die-out in nested sampling. The result is not analytic, but gives a simple rule of thumb, which I think is basically always fulfilled. So: do not worry about this!

    End-to-end proofs

    A limitation of current theoretical analyses is that most of these assume faithful likelihood-restricted prior sampling (LRPS). How can we go beyond this?

    Brendon Brewer solved this by embedding nested sampling within an MCMC framework. Fine, but most people do not use nested sampling that way.

    For nested sampling with step samplers, a key paper is “Unbiased and Consistent Nested Sampling via Sequential Monte Carlo” by Queensland University of Technology researchers Robert Salomone, Leah South et al., which present a proof by placing nested sampling within the sequential nested sampling umbrella of well-built-out theoretical analysis. There is a hole in that paper, which they are clear about: In SMC, there is a kernel which refreshes all points, while in NS, only one specific point is replaced. And so the connection is as of yet almost, but not quite, there.

    For nested sampling with region samplers, I just put out a theory paper analysing the replacement sampling space with Binomial point processes, which is the exact right tool for the job. This analysis is on the path to a end-to-end convergence proof of nested sampling, or at least bounding the errors.

    The nice aspect of convergence analyses is that once you have them, you know what you can tune. It does not work the other way around …

    By my recent count, there are over 6000 citations to the top 4 nested samplng packages. As machine learning becomes more prevalent, nested sampling analyses are used as the ground truth to compare against. Both of these facts should be justification to thoroughly understand nested sampling. Now we just need astronomers, statisticians and computer scientists to allocate funding for this research.

  • The state of nested sampling research in 2026: I) Scaling nested sampling to high-dimensions

    Since I wrote a systematic literature review (arxiv, scholar) in 2023, and written or contributed to several nested sampling packages, I’m someone who people turn to and ask, “what is next in nested sampling research”?

    In this first part, I will write about scaling nested sampling to high-dimensions.

    Contrary to conventional wisdom, nested sampling does not have poor scaling with dimensionality at all. If the prior equals the posterior, i.e., no information gain, the algorithm terminates successfully immediately, with high effective sample size.

    Two difficulties occur with the high information common with large datasets and hierarchical Bayesian models, where we learn an enormous amount of information. This can be easily seen with prior predictive checks – every prior realisation is plausible, but the observations are each highly peculiar in their own way, requiring a lot of fine tuning of the parameter space to find the posterior. This implies many nested sampling iterations, because in nested sampling, the number of iterations times the number of live points is proportional to the information gain: sqrt(N*K)=H

    The second difficulty is that the likelihood-restricted prior sampling used may scale poorly with dimensionality. The best technique so far is hit-and-run Monte Carlo (HARM) – miscalled slice sampling – first proposed used T Jasa & N Xiang – but introduced in detail by Will Handley’s PolyChord. Slice sampling is somewhat of a misnomer because we are not sampling the slice height, and in classical slice sampling one samples one dimension at a time. In any case, such MCMC approaches scale with dimensionality in a theoretically simple but practically difficult-to-understand way. One needs to tune M, the number of MCMC steps until the next point is decorrelated from the starting location (a random live point). Realistically the minimum M may be different throughout the run, but the worst-case will dominate and thus sets M which sets the compute budget. Annoying.

    Here are the current approaches:

    1. Make the run code faster – this is engineering: Some hacks here include vectorization, analytic marginalisation of variables not explicitly needed, and changing programming language. 2 research groups work on jax-based nested sampling implementations and parallelising. From an algorithm point of view, there is nothing here. Good marketing though …
    2. Focus on convex likelihoods: proximal nested sampling, ~1 researcher focused on imaging. This is an interesting approach not taken up by other research groups, yet.
    3. Give up on doing it right the first time and anneal towards the answer.
      • Diffusive nested sampling frames NS sampling as a MCMC process, with 1 active researcher (Brendon Brewer). The range of proposals within Diffusive Nested Sampling has not yet been fully explored, and a jax-based implementation would be good. I’ve had no luck getting the algorithm to behave though.
      • Snowballing nested sampling: Instead of expanding the number of steps until you get a nicely behaved run, expand the number of live points with a fixed number of steps.

    In my opinion, one of the most powerful techniques for making nested sampling run fast with many parameters is prior predictive checks. Often, the prior is misspecified, and information gain is unnecessarily high. Start from a reasonable prior.

    In high dimensions, you should also ask yourself whether the parameters have a meaning. If you are dealing with millions of pixel parameters and morphology, what are you doing? Do you really want to do Bayesian model comparison or do you just not know any other model comparison methods that may be a better fit?

    But what can you do. “High-dimensional inference” sounds cool and important, but I believe other inference axes are also worth pursuing, such as high information gain inference, phase transition inference, multi-modal inference.

    Because of the change in nature of the parameters, I don’t think nested sampling is the ideal algorithm for high-dimensional inference. My go-to-methods are

    a) divide and conquer importance resampling (as for example in PosteriorStacker). This is suitable for hierarchical models with low-dimensional per-object parameter spaces with medium to high information-gain per object. The parent distribution parameter space can be high dimensional.

    b) dynamic HMC. I’ve seen several people die on the hill of getting Stan to fit realistic astronomy models though. Differentiable programming with instrument responses is tough.

    The aversion of statisticians to nested sampling because they prefer SMC and the aversion of astrophysicists to SMC is a bit funny, given how close they are among the family of Monte Carlo algorithms. SMC requires some tuning of the tempering schedule, but has the advantage that one can potentially skip ahead to the posterior bulk. This is desirable for impatient researchers, such as those working on Gravitational Wave detections. A recent interesting paper was SMC+NUTS by Demasi+26.

    see also part II.

  • How many people does it take for society to be against you?

    Let’s say society is divided into two factions. One faction is univocally for equality between men and women in all areas. Another faction thinks that due to various factors, women and men should in some areas have different societal roles.

    People treat others according to their values.

    Let’s say the ratio is 80:20 in favor of your value system. You are a woman and interact with 100 people per day from all factions.

    What would this feel like?

    First, in interactions you would regularly interactions get comments that are inappropriate for your value system. Proportionally, you would get “Why are you still working?” and “When are you going to start working again?”. So you will get confronted with the minority values that most of society disagrees with, every day.

    In a previous post, I looked at the small number of bad actors it takes to make a community of neutral people a toxic environment.

    Most people you are friends with are probably sharing your values, or are neutral. So you will feel the society at large is wrong and against you. Even if you are in the majority.

    Major news outlets, because the 80% are boring non-news, will report on the incidents highlighting the 20% faction, skewing perception further.

    If the confidence of the 20% in expressing their values changes up and down over the years, it will look like major changes in society.

    Academia

    Now let’s consider the impact in academia. Much has been said about the leaky pipeline, which has many causes, including poor support and pressure on mothers, harassment, sexual misconduct, and unequal share of care work, among others.

    Still, let’s say in your institute, out of 10 women who enter, 8 are for equality and 2 are in favor of traditional gender roles.

    Now let’s say one of the 8 writes an email to the department in favor of equality, the other 7 may not feel the need to say anything because of course they agree. But the other 2 may give opposing views. They may also be supported by their equivalent camp of men. And now you have a feeling of being in a minority, and questioning whether you are right, why is everyone against you, can you fit in, etc.

    This skewing effect is probably most drastic on social media, where responding takes effort, and that effort is most commonly taken up for a backlash reaction, rarely for a supportive action.

    The skew can be addressed by stating your values openly, and by expressing support more often to people when they are so obviously right that it feels like it does not need to be said.

  • Simulation-based inference for X-ray astronomy

    In simulation-based inference, you train a neural network with pairs of model parameters θ and correspondingly data D, (θ, D). In the case of NRE, the network is a classifier, which predicts the probability that the data came from the (θ, D) stream corresponding to the joint p(θ,D) vs. a scrambled-up (θ, D) stream, corresponding to the respective priors p(θ)*p(D). With some algebra, this gives you p(D|θ), i.e., the likelihood. In the case of NPE, the network predicts p(θ|D) using a normalising flow, by optimizing the probability function to output most probability weight where θ is for the given D.

    The stream of (θ, D) is created by defining a prior p(θ), and a generator function D|θ which simulates data. Easy. Standard “forward-folding” (astronomer-speak) or generative modelling (statistician-speak).

    This scheme works well in a ML setting, where you have many tuples (θ, D), ideally, billions. Yes, you need a lot of data. This training cost hinders iterative model building.

    You can reuse standard industry network architectures to handle image data easily.

    Limitation 1: Varying observing conditions

    The scheme does not fit if you have variations in observing conditions. For example, if your instrument behaves slightly differently for each observation.

    Varying observing conditions are pervasive in X-ray astronomy because the object position on the focal plane changes the instrument response, both in the spectral and imaging sense. In optical astronomy, the equivalent is that you have different seeing conditions (filters are not considered time-variable). So we need to deal with this.

    [Aside on backgrounds: The background can also vary with position and time, which matters unless you are studying X-ray binaries and feel brave. The background can be simulated as well, with extending both θ and D, to background parameters and background region data, so this is not a deal-breaker.]

    The standard method then would be to either sample a random observing condition, and generate (θ, D), which would marginalise over the observing condition. Not wrong, but would lose information. The standard optical method is to derive the seeing conditions in the data analysis step from point sources in the image (e.g., SourceExtractor, photutils). This is reasonably fast. For X-ray astronomy, this is not possible, because the response is not location-invariant, so you cannot transplant information from nearby to the source.

    The proper way would be to sample not just p(θ, D), but p(θ, D | C) where C are the observing conditions, and define a conditional SBI (NRE or NPE) that gives its results conditional on an input C. If C was just the exposure time, and PSF FWHM, then we could make it part of D and sample it, and at inference time condition on it. However, in X-ray astronomy, C is a response matrix, and it is not trivial to smoothly sample from it, or to inform a SBI neural network about how to use it, because the SBI does not forward-fold. I suppose one could project C into a low-dimensional space.

    In any case, the samplings of C would likely be incomplete, and a new observation may not lie inside, erasing much of the cost savings at inference time that SBI should bring.

    Limitation 2: Poisson count data

    Why are there so few SBI – and, more generally – neural network papers with X-ray data? I think it is because neural networks suck with count data.

    You would think that you take a standard network that works with continuous data trained and knowing the L2 loss is proportional to a negative log-likelihood of a Gaussian switch for the loss for a negative log-likelihood of a Poisson and your done. But try to apply it to count data, and no, it doesn’t work. Why not? That’s a interesting ML research question. Correctly specified loss does not mean a trainable loss.

    So X-ray ML astronomers are hacking around by making X-ray count data Gaussian-like by smoothing, losing information in the process, or testing with trivial high-count situations.

    But once you bin finely enough in time or spectral channels (think XRISM, X-IFU), you are in the low count regime at least in those bins and have to deal with it.

    In my personal opinion, SBI is worth pursuing, but there are alternative approaches as well that scale to high dimensions that are not losing information. In any case, neural networks play a role, but it matters how you use them. Yes, I am being coy here on what is better, I’d like to be funded to resolve these issues.

  • Conflicts in politicial coalitions

    Coalitions become very unpopular when politicians are seen fighting and haggling. Calls for re-elections appear, together with the phrase “Instead of fighting all the time, the government should work.”

    Elections represent the choices of the people, ideally not just in the representatives but the direction they intend to lead towards. If people were to already agree in majority, the outcome would be a single party leading the country in one direction.

    If there is no party with a majority, the people are delegating the finding of a direction. On election day, this indecision is yet to be resolved. It is the job of politicians to resolve it, through consensus, compromise, conflict or otherwise.

    So I see politicians of different parties fighting as the delegated resolution of the people’s conflict. They are undergoing the process.

    In different countries, this may be more or less open. Perhaps in the past, in socialist countries and in dictatorships, this happens behind closed doors. Therefore, it appears orderly.

    As transparency is increased, these conflicts become more visible. Maybe many people are uneasy with just the knowledge that there is a conflict ongoing. Or they may reject the style of resolving the conflict of interests. The ugliest forms are narcissistic politicians doing name calling and bullying. Alternatives, such as consensus finding, exist and may be preferable by conflict-avoiding people.

    I also dislike conflicts myself, but I don’t mind delegating them to politicians to sort out.