Showing posts with label model. Show all posts
Showing posts with label model. Show all posts

2016-08-02

infer the expansion of the Universe without distances

As I note in this blog post, it is possible to infer that the Universe is expanding, even if you have only velocities and no distances. The idea is that you would marginalize out all distances and the Hubble Constant, and do a likelihood test (or equivalent, like, say, cross-validation). The two hypotheses would be a gas of galaxies with finite velocity dispersion but a well-defined mean rest-frame velocity through which we are moving (at unknown speed) vs an expanding gas, again with the finite velocity dispersion and through which we are moving. I think the test is very easy to set up and the problem is very easy to solve, and would demonstrate that it only takes a handful of velocities (but very good sky coverage!) to demonstrate expansion.

2012-10-21

find LRG-LRG double redshifts

Vivi Tsalmantza and I have found many double redshift in the SDSS spectroscopy (a few examples are published here but we have many others) by modeling quasars and galaxies with a data-driven model and then fitting new data with a mixture of two things at different redshifts. We have found that finding such things is straightforward. We have also found that among all galaxies, luminous red galaxies are the easiest to model (that's no breakthrough; it has been known for a long time).

Put these two ideas together and what have you got? An incredibly simple way to find double-redshifts of massive galaxies in spectroscopy. And the objects you find would be interesting: Rarely have double redshifts been found without emission lines (LRG spectra are almost purely stellar with no nebular lines), and because the LRGs sometimes host radio sources you might even get a Hubble-constant-measuring golden lens. For someone who knows what a spectrum is, this project is one week of coding and three weeks of CPU crushing. For someone who doesn't, it is a great learning project. If you get started, email me, because I would love to facilitate this one! I will happily provide consultation and CPU time.

2012-09-12

impute missing data in spectra

Let me say at the outset that I don't think that imputing missing data is a good idea in general. However, missing-data imputation is a form of cross-validation that provides a very good test of models or methods. My suggestion would be to take a large number of spectra (say stars or galaxies in SDSS), censor patches (multi-pixel segments) of them randomly, saving the censored patches. Build data-driven models using the uncensored data by means of PCA, HMF, mixture-of-Gaussians EM, and XD, at different levels of complexity (different numbers of components). Compare in their ability to reconstruct the censored data. Then use the best of the methods as your spectral models for, for example, redshift identification! Now that I type that I realize the best target data are the LRGs in SDSS-III BOSS, where the (low) redshift failure rate could be pushed lower with a better model. Advanced goal: Go hierarchical and infer/understand priors too.

2012-09-11

galaxy photometric redshifts with XD

Data-driven models tend to be very naive about noise. Jo Bovy (IAS) built a great data-driven model of the quasar population that makes use of our highly vetted photometric noise model, to produce the best-performing photometric redshift system for quasars (that I know). This has been a great success of Bovy's extreme deconvolution (XD) hierarchical distribution modeling code. Let's do this again but for galaxies!

We know more about galaxies than we do quasars—so maybe a data-driven model doesn't make much sense—but we also know that data-driven models (even ones that don't take account of the noise) perform comparably well to theory-driven models, when it comes to galaxy photometric redshift prediction. So a data-driven model that takes account of the noise might kick ass. This was strongly recommended to me by Emmanuel Bertin (IAP). In other news, Bernhard Schölkopf (MPI-IS) opined to me that it might be the causal nature of the XD model that makes it so effective. I guess that's a non-sequitur.

2012-09-04

find catastrophes in the stellar distribution

In Zolotov et al (2011) we asked the question: Might tiny dwarf galaxy Willman 1 be just a cusp in the stellar distribution of the Milky Way? If you generically have lines and sheets in phase space—and we very strongly believe that the Milky Way does—then generically you will have folds in those (in non-trivial projections they are required), and those folds generically produce catastrophes (localized regions of very high density) of various kinds (folds, cusps, swallowtails, and so on), which could mimic gravitationally bound or recently disrupted overdensities in the stellar distribution. The cool thing is that the catastrophes have quantitative two-dimensional morphologies that are very strongly constrained by mathematics (not just physics). The likelihood test we did in the Zolotov paper could easily be expanded into a search technique, maybe with some color-magnitude-diagram filtering mixed in. The catastrophes pretty much have to be there so get ready to get rich and famous! If you go there, send email to Scott Tremaine (IAS), who first proposed this idea to me.

2012-09-03

analyze quadratic star centroiding

Inside the core SDSS pipelines and inside the Astrometry.net source-detection code simplexy, centroiding—measurement of star positions in the image—is performed by fitting two-dimensional second-order polynomials to the central 3×3 pixel patch centered on the brightest pixel of each star. This is known to work far better than taking first moments of the light distribution (integrals of x and y times the brightness above background) for the (possibly obvious) reason that it is a quasi-justified (in terms of likelihood) fit.

Of course not all of the information about a star's position is contained in that central 3×3 pixel patch (and the method doesn't make use of any point-spread function information to boot). For this reason, Jo Bovy (IAS) and I did some work a few years ago to test it. Things intervened and we never finished, but our preliminary results were really surprising: For well-behaved point-spread functions, the two-dimensional quadratic fit in the 3×3 patch performed almost indistinguishably from fits that made use of the true point-spread function and larger patches. That is, it appeared in our early tests that the 3×3 patch does contain most of the centroiding information! A good research project would see how the 3×3 patch inference degrades relative to the point-spread-function inference, as a function of PSF properties and the signal-to-noise, with an eye to analyzing when we need to be thinking about doing better. I will call out Adrian Price-Whelan (Columbia) here, because he is all set up with the machinery to do this!

2012-09-01

remove satellite trails from arbitrary astronomical imaging

Satellite trails appear as long lines in astronomical imaging, often nearly unresolved or slightly resolved. They are easy to find, fit, and subtract away, at least in principle. I have had several undergraduate researchers, however, who got close but couldn't deliver a robust, reliable piece of code.

The code I imagine takes an image (and an optional inverse variance image). It identifies if the image contains a satellite trail (possibly using the Hough Transform and some heuristics). If it does, it fits the trail using robust fitting techniques. If that all works, it returns to the user an updated image and an updated inverse variance map. Not hard! The only hard parts are making it robust and making it fast. I have a lot of good ideas on both parts of that; I think this is very do-able, and it is only a few weeks work for the right person. It would be hella useful too, especially for the human-viewable image projects I am working on. Enhanced goal: Fit for satellite tumbling or blinking (both things are common in the data I have).

2012-08-26

what is the spectrum of dust attenuation?

The SDSS has taken spectra of thousands of F-type stars, at different distances and through different amounts of interstellar dust. These stars were chosen for calibration purposes; they were chosen because they have very well-understood and consistent spectra. These have been used to calibrate the SDSS telescope, but they can also be used to calibrate interstellar dust.

The general procedure would be to start by measuring the equivalent widths of a few absorption lines—preferably a couple of Balmer lines and a couple of metal lines— consistently for all F-stars. These line EWs would provide a dust-indpendent temperature and metallicity indicator for all the stars. Compare the spectra of the F-stars at different reddening but fixed absorption-line equivalent widths (and therefore fixed temperature and metallicity) to get the dust attenuation at resolution of a few thousand. There probably isn't anything interesting there, but if there is it would be a valuable discovery.

The easiest way to do this project is by spectral stacking, but there might be methods that build a non-linear model of the stellar spectrum with three controlling parameters: Balmer EW, metal EW, and SFD-dust-map amplitude. I started discussing this project many years ago with Karl Gordon (STScI); if you want to give it a shot, send us both email for ideas (if you want to; otherwise do it and surprise us!).