Like many astronomical missions and data sets, the NASA Kepler satellite imaging is crowded; there is no part of the imaging that contains reliably isolated stars. For this reason, it is hard to infer (from the data) the point-spread function of the instrument simply; any PSF inference requires modeling the images as a crowded image of many overlapping stars (of unknown brightnesses and positions). However, when a star is subject to a planetary or stellar transit, the change in the scene during the transit should be modeled well as a PSF-shaped deficit. So we should be able to infer the PSF from these deficits. The transits are rare, so they rarely overlap. They are also very faint (low-amplitude) events, but (a) the eclipsing-binary stellar transits are not so faint (and very common), and (b) Kepler has good signal-to-noise even on planetary transits.
2016-01-04
2015-01-10
video magnification on eta Carinae
We should apply the video magnification methods of Bill Freeman (MIT) to the spectacular HST images of expanding star eta Carinae. The motion in the HST images is a bit subtle, but it sure wouldn't be if we fired up the Freeman tricks. One of the cleverest of his tricks is to take the Fourier Transform of each image and then magnify the phase differences. This ensures that the thing being magnified is global and smooth, in some sense. The one respect in which Freeman's methods would have to be generalized is that the data are not from a single, video source: The different images were taken through different filters on different cameras and with different exposure times.
2014-07-10
predict spectra using photometry
The SDSS spectra can be thought of as "labels" for objects detected in the imaging, each of which has ugriz photometry and some shape and position parameters. Can we train a model with this enormous amount of data to predict the spectra using the photometry? One thing that says "yes" is that photometric redshifts (for galaxies and quasars), photometric distances (for stars), and photometric temperatures and metallicities (for stars) all work well. One thing that says "no" is that there is far more information (in a technical sense) in the spectra than in the photometry. All this said, it is an absolutely great "Data Science" demonstration project, and it might create some new ideas for LSST-era astrophysics projects. In principle, it will also get us predictions about the spectral types and redshifts of many objects that lack spectra!
2013-11-16
extract low-resolution spectra from diffraction spikes
In imaging from a telescope with a secondary on a spider (for example, in HST imaging), bright stars show diffraction spikes. More generally, the outer parts of the point-spread function are related to the Fourier Transform of the small-scale features in the entrance aperture. The scale at which this Fourier Transform imprints on the focal plane is linearly related to wavelength (just as the angular size of the diffraction-limited PSF goes as wavelength over aperture).
This means that the diffraction spikes coming from stars contain low-resolution spectra of those stars! That is, you ought to be able to extract spectral information from the spikes. It won't be good, but it should permit measurements of colors or temperatures or SED slopes with even single-band imaging, and aid in star–quasar classification. Indeed, in HST press-release images, you can see that the diffraction spikes are little "rainbows" (see below).
The project is to take wide-band imaging from HST, in fields where stars have been measured either in multiple bands or else spectroscopically, and show that some of the scientific results could have been extracted from the single, wide band directly using the diffraction features.
2012-11-25
limits to ground-based photometry?
A conversation with Nick Suntzeff (TAMU) in Lawrence, KS, brought up the great idea (Nick's, not mine) to figure out why ground-based photometry of stars never gets better than a few milli-mags in precision. Seriously people, Kepler is at the part-per-million or better level. Why can't we do the same from the ground? Why not at least part-per-hundred-thousand? Is it something about the scintillation, the transparency, the point-spread function, the detector temperature, scattered light, sky emission, sky lines, what? Not sure how to proceed, but the project could make the next generation of projects orders of magnitude less expensive. I guess I would start by taking images of a star field with many different (very different) exposure times and at different twilight levels (Suntzeff's idea again). Could it be that all we need is better software?
2012-10-24
find or rule out a periodic universe (via structure)
Questions from Kilian Walsh (NYU) today reminded me of an old, abandoned idea: Look for evidence of a periodic universe (topological non-triviality) in the large-scale structure of galaxies. Papers by Starkman (CWRU) and collaborators (one of several examples is here) claim to rule out most interesting topologies using the CMB alone. I don't doubt these papers but (a) they effectively make very strong predictions for the large-scale structure and (b) if CMB (or topology) theory is messed up, maybe the constraints are over-interpreted.
The idea would be to take pairs of finite patches of the observed large-scale structure and look to see if there are shifts, rotations, and linear amplifications (to account for growth and bias evolution) that make their long-wavelength (low-pass filtered) density fields match. Density field tracers include the LRGs, the Lyman-alpha forest, and quasars. You need to use (relatively) high-redshift tracers if you want to test conceivably relevant topologies.
Presumably all results would be negative; that's fine. But one nice side effect would be to find structures (for example clusters of galaxies) residing in very similar environments, and by similar
I mean in terms of full three dimensional structure, not just mean density on some scale. That could be useful for testing non-linear growth of structure.
2012-10-21
find LRG-LRG double redshifts
Vivi Tsalmantza and I have found many double redshift in the SDSS spectroscopy (a few examples are published here but we have many others) by modeling quasars and galaxies with a data-driven model and then fitting new data with a mixture of two things at different redshifts. We have found that finding such things is straightforward. We have also found that among all galaxies, luminous red galaxies are the easiest to model (that's no breakthrough; it has been known for a long time).
Put these two ideas together and what have you got? An incredibly simple way to find double-redshifts of massive galaxies in spectroscopy. And the objects you find would be interesting: Rarely have double redshifts been found without emission lines (LRG spectra are almost purely stellar with no nebular lines), and because the LRGs sometimes host radio sources you might even get a Hubble-constant-measuring golden lens
. For someone who knows what a spectrum is, this project is one week of coding and three weeks of CPU crushing. For someone who doesn't, it is a great learning project. If you get started, email me, because I would love to facilitate this one! I will happily provide consultation and CPU time.
2012-10-10
find or rule out ram pressure stripping in galaxy clusters
We know a lot about the scalar properties of galaxies as a function of clustocentric distance: Galaxies near cluster centers tend to be redder and older and more massive and more dense than galaxies far from cluster centers. We also know a lot about the tensor properties of galaxies as a function of clustocentric distance: Background galaxies tend to be tangentially sheared and galaxies in or near the cluster have some fairly well-studied but extremely weak alignment effects. What about vector properties?
Way back in the day, star NYU undergrad Alex Quintero (now at Scripps doing oceanography, I think) and I looked at the morphologies of galaxies as a function of clustocentric position, with the hopes of finding offsets between blue and red light (say) in the direction of the cluster center. These are generically predicted if ram-pressure stripping or any other pressure effects are acting in the cluster or infall-region environments. We developed some incredibly sensitive tests, found nothing, and failed to publish (yes I know, I know).
This is worth finishing and publishing, and I would be happy to share all our secrets. It would also be worth doing some theory or simulations or interrogating some existing simulations to see more precisely what is expected. I think you can probably rule out ram-pressure stripping as a generic influence on cluster members, although maybe the simulations would say you don't expect a thing. By the way, offsets between 21-cm and optical are even more interesting, because they are seen in some cases, and are more directly relevant to the question. However, it is a bit harder to assemble the unbiased data you need to perform a sensitive experiment.
2012-09-16
scientific reproducibility police
At coffee this morning, Christopher Stumm (Etsy), Dan Foreman-Mackey (NYU), and I worked up the following idea of Stumm's: Every week, on a blog or (I prefer) in a short arXiv-only white paper, one refereed paper is taken from the scientific literature and its results are reproduced, as well as possible, given the content of the paper and the available data. I expect almost every paper to fail (that is, not be reproducible), of course, because almost every paper contains proprietary code or data or else is too vague to specify what was done. The astronomical literature is particularly interesting for this because many papers are based on public data; for those it comes down only to code and procedures; indeed I remember Bob Hanisch (STScI) giving a talk at ADASS showing that it is very hard to reproduce the results of typical papers based on HST data, despite the fact that all the data and almost all the code people use on them are public.
Stumm, Foreman-Mackey, and I discussed economic models and incentive models to make this happen. I think whoever did this would succeed scientifically, if he or she did it well, both because it would have huge impact and because it would create many new insights. But on the other hand it would take significant guts and a hell of a lot of time. If you want to do it, sign me up as one of your reproducibility agents! I think anyone involved would learn a huge amount about the science (more than they learn about reproducibility). In the end, it is the community that would benefit most, though. Radical!
2012-09-12
impute missing data in spectra
Let me say at the outset that I don't think that imputing missing data is a good idea in general. However, missing-data imputation is a form of cross-validation that provides a very good test of models or methods. My suggestion would be to take a large number of spectra (say stars or galaxies in SDSS), censor patches (multi-pixel segments) of them randomly, saving the censored patches. Build data-driven models using the uncensored data by means of PCA, HMF, mixture-of-Gaussians EM, and XD, at different levels of complexity (different numbers of components). Compare in their ability to reconstruct the censored data. Then use the best of the methods as your spectral models for, for example, redshift identification! Now that I type that I realize the best target data are the LRGs in SDSS-III BOSS, where the (low) redshift failure rate could be pushed lower with a better model. Advanced goal: Go hierarchical and infer/understand priors too.
2012-09-11
galaxy photometric redshifts with XD
Data-driven models tend to be very naive about noise. Jo Bovy (IAS) built a great data-driven model of the quasar population that makes use of our highly vetted photometric noise model, to produce the best-performing photometric redshift system for quasars (that I know). This has been a great success of Bovy's extreme deconvolution (XD) hierarchical distribution modeling code. Let's do this again but for galaxies!
We know more about galaxies than we do quasars—so maybe a data-driven model doesn't make much sense—but we also know that data-driven models (even ones that don't take account of the noise) perform comparably well to theory-driven models, when it comes to galaxy photometric redshift prediction. So a data-driven model that takes account of the noise might kick ass. This was strongly recommended to me by Emmanuel Bertin (IAP). In other news, Bernhard Schölkopf (MPI-IS) opined to me that it might be the causal nature of the XD model that makes it so effective. I guess that's a non-sequitur.
2012-09-05
show that low-luminosity early-type galaxies are oblate
Here's an old one from the vault: Plot the surface brightness of early-type galaxies (red, dead) as a function of ellipticity and show that surface brightness rises with ellipticity. This is what is expected if early-type galaxies are transparent and oblate. I know from nearly completing this project many years ago that this will work well for lower-luminosity early types and badly for higher-luminosity early types. The cool thing is that, under the oblate assumption, the true three-dimensional axis-ratio and three-dimensional central stellar density distribution function can be inferred from the observed two-dimensional distributions under the (weak) assumption of isotropy of the observations. That assumption isn't perfectly true but it is close. You can use high signal-to-noise imaging and SDSS spectroscopy to do the object selection, so observational noise in selection and measurement won't provide big problems.
This is another Scott Tremaine (IAS) project. Mike Blanton (NYU) and I basically did this many years ago with SDSS data, but we never took it through the last mile to publication, so it is wide open. Actually, it seems likely that someone has done this previously, so start with a literature search! Bonus points: Figure out what's up with the high-luminosity early types. They are either triaxial or a mix of oblate and prolate.
2012-09-04
find catastrophes in the stellar distribution
In Zolotov et al (2011) we asked the question: Might tiny dwarf galaxy Willman 1 be just a cusp in the stellar distribution of the Milky Way? If you generically have lines and sheets in phase space—and we very strongly believe that the Milky Way does—then generically you will have folds in those (in non-trivial projections they are required), and those folds generically produce catastrophes (localized regions of very high density) of various kinds (folds, cusps, swallowtails, and so on), which could mimic gravitationally bound or recently disrupted overdensities in the stellar distribution. The cool thing is that the catastrophes have quantitative two-dimensional morphologies that are very strongly constrained by mathematics (not just physics). The likelihood test we did in the Zolotov paper could easily be expanded into a search technique, maybe with some color-magnitude-diagram filtering mixed in. The catastrophes pretty much have to be there so get ready to get rich and famous! If you go there
, send email to Scott Tremaine (IAS), who first proposed this idea to me.
2012-08-29
build a paragraph-level index for arXiv
There is plenty of technology ready-to-use that does author-topic modeling, in part because so many of the machine-learning community's successes have been in text processing and classification. Some of this technology has been run on the arXiv at the abstract, title, meta-data level, but rarely at the full-text level. If we ran it there, we could build an author-topic model over every paragraph in the full corpus of the arXiv. This would tell us what every paragraph is about (and, amusingly, who wrote every paragraph), with (if we do it right) probabilistic output. That is, it would give probability distributions over what every paragraph is about.
If done correctly, and with a little hand-labeling of classes (and this is easy because each class would have characteristic words and phrases and authors), this could lead to a complete and exceedingly useful paragraph-level index into the arXiv. But even without the hand-labeling it would be incredibly useful: It could be used to find paragraphs in the literature that are useful and relevant to every paragraph you have written in one of your own papers, thus locating related work you might not know about. It could drive search services that find paragraphs relevant to your search terms even when they don't, themselves, contain those terms. And so on!
Dan Foreman-Mackey (NYU) first made it clear to me that this would be possible and he also took some steps towards making it happen. David Blei (Princeton) suggested to me that even from a machine-learning perspective the outcomes could be very interesting. The plagiarism paper from the arXiv people suggests working at PDF level rather than LaTeX source level; I am not sure myself which to do.
2012-08-28
will chemical tagging work?
Chemical tagging is the name given to the idea that we can match up stars in abundance space (detailed chemical properties) as well as kinematic space to figure out the origin and common orbits of stars in the Milky Way. Because it would be so valuable to figure out that different stars shared a common origin at formation (for things like orbit inference), chemical tagging could enormously improve the precision of any dynamical or galaxy-formation information coming from next-generation surveys.
In the many conversations I have seen about chemical tagging, arguments break out about whether it is possible to measure the chemical abundances of stars of different temperatures and surface gravities comparably. That is: Can we figure out that this F star has the same abundances as this other K star? Or this red giant and this main-sequence star? And it is certainly not clear: Chemical abundances are not measured at enormous precision and there are many possible biases, sources of variance, and systematic error.
My proposal is that we ask these questions not in the space of the outputs of chemical-abundance models but rather in the space of stellar spectroscopy observables. The question becomes not are the models good enough?
but rather is there information in the data?
And there needs to be information sufficient to distinguish thousands (yes that is the goal) of chemically distinct sub-populations.
If there is sufficient information, then in the dozens-to-hundreds-of-dimensions space of all possible absorption-line measurements (plus stellar temperature), do we see thousands of distinct families of (possibly very complex) one-dimensional loci (each locus being a birth-mass-sequence at fixed chemical abundances and age)? The idea would be to do this purely in the space of spectra but—probably necessarily—relying heavily on models to guide the eye
(or really guide the code) where to look.
I have discussed this with Ken Freeman (ANU) and Mike Blanton (NYU), but as far as I know, no-one is working on it. Blanton had the great idea that we don't really need to make spectral features before starting. The question does the distribution of stellar spectra split up into many tiny, thin, curvy lines in spectrum space?
can be asked with just well-calibrated spectra. And we have lots of those!
2012-08-27
emission-line clustering and classification
The BPT diagram has been incredibly productive in classifying galaxies into star-forming and AGN-powered classes. However, the diagram only shows two ratios of nearby lines; ratios of nearby lines so that dust and spectrograph calibration don't mess up the data, and two because it is a single two-dimensional plot. There might be many features in emission-line space sitting undiscovered in the data; there might be many sub-classes and rich structure within the star-forming and AGN groups.
From a data perspective, times have really changed since BPT: (1) There are dozens (well, a dozen) of visible lines in hundreds of thousands of spectra. (2) We have good noise models for the line measurements and this is especially important when they get low in signal-to-noise (as they do if you want to use many lines. (3) We have very well-calibrated spectra now, even spectrophotometrically good to a few percent in the SDSS. (4) The effects of dust attenuation are pretty well understood in the optical. So let's go high dimensional and find all the complex structure that must be there!
The first step is to measure all the lines in a long list, and measure them even when the signal-to-noise is low. We don't care about detections we care about measurements with well-understood noise. The second step is to develop dust-insensitive metrics: What is the distance
in data space between two sets of dust-line measurements as a function of noise but marginalizing out the dust affecting each spectrum? Now in that space, let's do some clustering.
I have done nothing on this except discuss it, years ago, with John Moustakas (Siena College). At that time, we were thinking in terms of generating archetypes with an integer program (with my now-deceased guru Sam Roweis). You could use things like support vector machines (great for these kinds of tasks) but we have no labels to classify on. The idea is to find classes not yet discovered! Also SVMs are not sensitive to the uncertainties in the data. I would recommend something like extreme deconvolution which does density estimation of the noise-deconvolved distribution. It can deal with very low signal-to-noise data gracefully. It would have to be modified, however, to project out (marginalize out) the dust-extinction direction in line space. Not impossible but not trivial either.
2012-08-26
what is the spectrum of dust attenuation?
The SDSS has taken spectra of thousands of F-type stars, at different distances and through different amounts of interstellar dust. These stars were chosen for calibration purposes; they were chosen because they have very well-understood and consistent spectra. These have been used to calibrate the SDSS telescope, but they can also be used to calibrate interstellar dust.
The general procedure would be to start by measuring the equivalent widths of a few absorption lines—preferably a couple of Balmer lines and a couple of metal lines— consistently for all F-stars. These line EWs would provide a dust-indpendent temperature and metallicity indicator for all the stars. Compare the spectra of the F-stars at different reddening but fixed absorption-line equivalent widths (and therefore fixed temperature and metallicity) to get the dust attenuation at resolution of a few thousand. There probably isn't anything interesting there, but if there is it would be a valuable discovery.
The easiest way to do this project is by spectral stacking, but there might be methods that build a non-linear model of the stellar spectrum with three controlling parameters: Balmer EW, metal EW, and SFD-dust-map amplitude. I started discussing this project many years ago with Karl Gordon (STScI); if you want to give it a shot, send us both email for ideas (if you want to; otherwise do it and surprise us!).
2012-08-24
get SDSS colors and magnitudes for very bright stars
The SDSS saturates around 14th magnitude. However, (a) the gains are set such that the CCD pixels saturate before the analog-to-digial read-out saturates, and (b) the bleeding of charge on the CCD is essentially charge-conserving. Also, when very bright stars cross the readout register in the CCD, they leave a thin 2048-pixel line across the full camera column. And also also, the stars have well-defined diffraction spikes that are visible to large angular radii.
No-one says this is easy; this is a blog of good ideas
not easy ideas
: For one, the detector may become weakly nonlinear shortly before CCD pixel saturation; that is, the effective gain may be lower at brighter magnitudes; any project would have to look carefully into this, and you don't have variable exposure times to use (all of SDSS was taken at 55-second exposure time for very important reasons). For another, the shape and size of the diffraction spikes might be a strong function of position in the focal plane. However, I have hope, because the charge bleeds are so very very beautiful when inspected in detail.
Some prior work on this has been done by myself and Doug Finkbeiner (Harvard). It would be worth checking in with Fink before embarking.
2012-08-23
bimodality search or kurtosis components analysis
Take the SDSS spectra (which are beautifully calibrated spectrophotometrically) and interpolate them onto a common rest-frame (de-redshifted) wavelength grid. Do clever things to interpolate over missing and corrupted data where necessary; this might involve performing a PCA and using the PCA to patch and then re-doing PCA and so on. Then re-normalize the data so that the amplitudes
of all the spectra are the same; I am being vague here because I don't know the best choice for definition of amplitude
. This is all pre-conditioning for the data; in principle the recommendation here could be applied to any data set; I am just proposing the SDSS spectra.
Now search for a unit-norm (or otherwised normalized) eigenspectrum such that when you dot all pre-conditioned SDSS spectra onto the eigenspectrum, you obtain a distribution of coefficients (dot products) that has minimum kurtosis. That is, instead of finding the principal components—the components with maximum variance—we will look for the platykurtic components—the components with minimum kurtosis. If you are stoked, search the orthogonal subspace for the next-to-minimum kurtosis direction and so on.
Why, you ask? Because low-kurtosis distributions are bi-modal. Indeed, early experiments (performed by Vivi Tsalmantza (MPIA) and myself back in 2008) indicate that this will identify the eigenspectra that best separate the red sequence
galaxies from the blue cloud
. If you really want to go to town invent a bimodality scalar that is better than kurtosis.
One note: Optimization is a challenge. This sure ain't convex. My approach back in the day was to throw down randomly generated spectra, choose ones that happened to hit fairly low kurtosis, and optimize locally from those.
2012-08-22
cosmic-ray identification
Take a set of HST data from one filter and exposure time (to start; later we will generalize) that have been CR-split (meaning: two images at each pointing). Shift-and-difference these split images to confidently identify a large number of cosmic rays. Pull out 5x5 image patches centered on cosmic-ray-corrupted pixels and 5x5 image patches not centered on cosmic-ray-corrupted pixels. Use these labeled data as training data for a supervised method that finds cosmic rays in single-image (not-CR-split) data. Improve value of HST data for all and obtain enormous financial gift from NASA in thanks (well, not really).
Notes: The most informative pixel patches will be those with faint cosmic ray pixels and those with bright stars that mimic cosmic rays. Some of this work has been started with (now graduated) NYU undergraduate Andrew Flockhart.