New Methods in Theory & Applications of Nonlinear Filtering by Andrew Papanicolaou B. S., University of California at Santa Barbara 2003 M. S., University of Southern California 2007 A Dissertation submitted in partial fulfillment of the requirements for the Degree of Doctor of Philosophy in the Division of Applied Mathematics at Brown University Providence, Rhode Island May 2010 c Copyright 2010 by Andrew Papanicolaou This dissertation by Andrew Papanicolaou is accepted in its present form by the Division of Applied Mathematics as satisfying the dissertation requirement for the degree of Doctor of Philosophy. Date Boris Rozovsky, Director Recommended to the Graduate Council Date Constantine Dafermos, Reader Date George Karniadakis, Reader Approved by the Graduate Council Date Sheila Bonde, Dean of the Graduate School iii Vitæ Andrew Papanicolaou was born on 1st of January of 1980. As a youth he enjoyed jazz music and the study of its complex rhythms and harmonies. In the spring of 2003, he received a bachelors degree in mathematical sciences from the University of California at Santa Barbara, and matriculated at the University of Southern Califor- nia shortly thereafter to work on a masters degree in financial mathematics. In the fall of 2007, with his masters degree completed, he arrived at Brown University to write his PhD thesis on nonlinear filtering with Boris Rozovsky. iv “With Probability One” v Dedicated to My Parents vi Acknowledgements A very special thanks to my advisor for his mentorship over the past six years both at USC and here at Brown. For all my achievements thus far, he is the person to whom I give all the credit. Without his support I would never have made it this far. Special thanks to Michael Black at Brown University for introducing me to his work on human tracking with Bayesian methods. A very special thanks to the late Professor David Gottlieb whose profound wisdom in his last year encouraged me to work hard. vii Preface I certainly did not plan to earn a doctorate of applied mathematics when I started graduate school in the Fall of 2004. I enrolled in the financial mathematics program at USC with intentions of starting a career in banking as soon as I completed a Master’s degree. However, things never quite work out the way you plan, and now after six years of studies I am at Brown finishing my PhD thesis on nonlinear filtering. I am happy to be finishing and will look forward to my future work as a post-doc. Andrew Papanicolaou Providence, Rhode Island viii Contents Vitæ iv Dedication vi Acknowledgements vii Preface viii List of Tables xii List of Figures xiii Chapter 1. Introduction 1 Chapter 2. Filtering for Stochastic Volatility Models & Volatility Swap Contracts 6 1. Historical Volatility, Variance-Swaps, and the VIX 8 1.1. ‘Fair’ Valuation of Volatility Swaps 9 1.2. The VIX 11 2. Filtering and Smoothing Estimates of Stochastic Volatility 13 2.1. Filtering 14 2.2. Smoothing 16 2.3. Parameter Re-Estimation 17 2.4. Comparison of Spot-Estimates 20 3. Comparative Analysis of Filtering to VIX 21 3.1. Filtering and Smoothing Estimates of Averaged Volatility 21 3.2. Filtering and Smoothing Volatility Swaps 24 3.3. Statistical Fitting of Swap Returns as a function of the VIX 25 4. Chapter Summary 27 ix Chapter 3. Tracking Intentions from Streaming Video: A Double HMM Approach 29 1. Models and Algorithms for Tracking Intentions from Image Data 32 1.1. Forward Baum-Welch Equation 37 1.2. A Tracking Intentions Example 38 2. Filtering Techniques for Image Data Applications 41 2.1. Scaling Procedure for Small Likelihoods 41 2.2. Filtering With Limited Computing Resources 44 2.3. Banks of Filters for Parallel Computing 47 2.4. Particle Filtering Algorithm 48 2.5. Sampling on a High-Dimensional State-Space 49 3. Numerical Simulation 50 3.1. Pinocchio Example 51 3.2. Filtering Algorithm for Pinocchio 54 3.3. Simulations and Results 57 4. Chapter Summary 62 Chapter 4. A Kramers-Smoluchowski Approximation Applied to a Multi-Scale HMM 64 1. Target Tracking Models with Fast Mean-Reversion 67 1.1. Observations & Maximum Likelihood 70 1.2. The Non-Linear Filter 72 1.3. Fast Mean Reversion 73 1.4. FMR Models & Baum-Welch Equations of Reduced Dimension 74 2. Analysis of the Generalized Kramers-Smoluchowski Approximation 75 2.1. The Generalized Kramers-Smoluchowski Approximation 75 2.2. Kramers-Smoluchowski Approximation in Discrete Time 79 2.3. Numerical Simulations for the Kramers-Smoluchowski Approximation 80 2.4. Simulation of FMR Filtering 84 3. Chapter Summary 91 x Chapter 5. Nonlinear Filters of Reduced Dimension for Multi-Scale Stochastic Models 92 1. Definitions, Parameters & the MainResult 93 1.1. Fast Mean-Reversion and a Fast Filter 96 2. Tightness of Marginal Measures and Weak Convergence to a Unique Limit 99 2.1. Tightness 100 3. Numerical Simulations 106 3.1. Discrete Time Schemes 106 3.2. Error Graphs and Plots 108 3.3. Empirical Estimates of Optimal Sampling Rate 109 4. Chapter Summary 110 Chapter 6. Conclusion 112 Appendix A. Double Markov Chains 113 Bibliography 118 xi List of Tables 1 Tracking Intentions Example 39 2 Scaled Nonlinear Filtering Algorithm 45 3 Bank of Particle Filters 49 xii List of Figures 1 The S&P 500 from 1990 to 2004 and the VIX during the same time. 12 q 2 Given the S&P 500 data from 1990 to 2004, this plot show the filtering estimate E[σt2k |Fk ] compared with η · V IXtk where η is the standard deviation from table (8). Although the VIX is not a true spot-estimate of volatility, it is still worthwhile to observe how it relates to filtering. 15 q 3 Given the S&P 500 data from 1990 to 2004, this plot shows the filtering estimate of E[σt2k |Fk ] with parameters re-estimated by the EM algorithm, compared to η · V IXtk where η is the standard deviation from table (8). 19 4 Estimates (18) and (19) compared to η · V IX. The parameter η ≈ .73 is the empirical standard deviation of the recovered noise under the VIX data, as was shown in tables (8) and (17). 23 5 Top Left: RV[0,T ] plotted as a function of V IX. Top Right: RV[0,T ] for the current month plotted as a function of RV[0,T ] for the previous month, which shows the returns on the mark-to-market swap contracts. Bottom Left: SV plotted as a function of the V IX with Λ estimated by the jump frequencies of the VIX. Bottom Right: F V plotted as a function of the V IX also with Λ estimated by the jump frequencies of the VIX. 26 6 Top Left: RV[0,T ] plotted as a function of V IX. Top Right: RV[0,T ] for the current month plotted as a function of RV[0,T ] for the previous month. Bottom Left: SV plotted as a function of the V IX with Λ re-estimated by EM. Bottom Right: F V plotted as a function of the V IX also with Λ re-estimated by EM. 28 1 The flow diagram for a bank of nonlinear filters. At each time step the input is the previous bank of posterior distributions and the latest observation, from which a bank of updated posteriors is returned. 46 2 The only street lamp in the city which Pinocchio may or may not intend to shoot out. 51 3 The alphabet of poses which Pinocchio can take. 52 xiii 4 We see here how the spatial location is a mapping of a pose from the center of the street lamp scene. The Pinocchio in the center illustrates H(i, 0), and each arrow represents a possible value of Xk and how it will spatially relocate Pinocchio within the scene. 53 5 Low-noise observations of Pinocchio walking around near the street lamp. 54 6 A zero-noise example: The log-likelihood of the data for a realization {θk1 , Xk }100 k=0 generated under Γ{11} when γ = 0 is simply the log-probability of the realization, {ij} {11} {ij} Lk = log P((θ1 , X)|θ0 = (i, j)). The log-probability L100 is significantly larger than L100 for (i, j) 6= (1, 1). 59 7 High Noise Example. This figure shows a high noise example of Pinocchio when he has a gun and his intentions are hostile. The pictures are the snapshots taken by the surveillance camera at times k = 1, 22, 44, 65, and although this realization of Pinocchio would have shot out the light, the decision to arrest Pinocchio was made with plenty of time to spare. 60 8 High Noise Example. This figure shows the lag-weighted posterior statistic Sk as described in Section 3.2. It is plotted against time with a = .9 and α = .95. In this particular realization of Pinocchio, he shoots out the light at time k = 59, but when the ATR&D system is working it will arrest him earlier at time k = 32. 60 9 Unsuccessful Tracking Example. Suppose that we a have a realization (θ1 , X) = {θk1 , Xk }100 k=1 generated under Λ{11} . The figure plots the probability of the HMM (θ1 , X) under all log Γ{01} (θk1 , Xk |θk−1 1 P four possible probability laws, from which we see that k , Xk−1 ) ≈ log Γ{11} (θk1 , Xk |θk−1 1 P k , Xk−1 ). This means that the ATR&D may be confident that Pinocchio has a gun, but it will have difficulty in deciding whether his behavior is hostile or benign. 62 10 Unsuccessful Tracking Example. The lag-weighted posteriors, Sk . When Λ{11} and Λ{10} are very similar it makes it hard for the ATR&D to make a decision. We see in this figure how Sk never exceeds the 95% threshold. 62 1 A sample of the process (θt , Vta , Xta ) in two dimensions with the horizontal axis being time. The mean reversion rate a is such that Vta reverts to its mean fairly quickly with time. 72 2 This figure displays the convergence of the process Xka to the Kramers-Smoluchowski approximation as a → ∞. The sample processes are generated using equations (44), (45) and xiv (46). For several realizations of (θk , Bk ) the exact solution is compared to the approximation. q P The error is the empirical average of MSE, M SE = E k |Xka − X ˜ k |2 where X ˜ k denotes the Kramers-Smolushowski approximation. 82 3 In this figure the log-error of Xka compared with its Kramers-Smoluchowski approximation is computed as a → ∞. The sample processes are generated using equations (44), (45) and (46). „q « The vertical axis in the plot is an empirical approximation of log ˜ k |2 where X E k |Xka − X P ˜k denotes the Kramers-Smolushowski approximation. 83 4 An example of an observed image from which the filter must infer the location and state of the target. The dim glow is the true location of the target. 85 5 The FMR reduced filter from equation (39) with a = 250 and ∆t = .01. Both the processes θk and Xka are tracked based on the information provided by a sequence of frames that look like the one given in figure 4. The significant spots are the changes in the kinematic mode. The filter is able to catch the changes without tracking the process Vka . 86 6 The filtering of the processes as seen from the camera’s view-finder. The red ‘X’ marks the target’s true location, and the blue ‘O’ marks the filtering estimate of the location. 87 7 The FMR reduced filter from equation (39) with a = 40 and ∆t = .01. There is significant error in tracking Θk because the (54) is not valid for this parameterization of the model. 89 8 As a decreases, the error increases. This plot shows the proportion of time that the MAP estimator of the reduced filter is wrong as a function of a. The vertical axis is the error, err = E|θk − θˆkmap |0 = E1θˆmap 6=θk . 90 k 1 The error of the averaged filter decreases to the that of the optimal filter as  & 0. Also, the signal-to-noise ratio decreases with , and so the error of the optimal filters decreases too. 108 2 The penalized error term, E  + 21 ∆t2 and E  + log(1 + ∆t), for different ∆t and . Notice how the optimal ∆t increases with . 111 xv Abstract of “New Methods in Theory & Applications of Nonlinear Filtering” by Andrew Papanicolaou, Ph.D., Brown University, May 2010 Hidden Markov models are used in countless signal processing problems, and the associated nonlinear filtering algorithms are used to obtain posterior distributions for the hidden states. One reason why posterior distributions are so important is because they are used to compute estimates which are optimal given a history of observed data. However, it is often the case that implementation of these algorithms is near impossible because of the curse-of-dimensionality which results from testing every possible hypothesis. This thesis explores new applications of nonlinear filtering and addresses several issues in algorithm implementation. CHAPTER 1 Introduction The concentration of this thesis is the various applications of nonlinear filtering and the associated challenges in algorithm implementation. There are four main chapters in this dissertation, each of which presents a self-contained analysis of a problem that is either a new and novel applications of nonlinear filtering, or a new method for deriving fast-filtering algorithms for hidden Markov models with multiple time-scales. A nonlinear filtering algorithm is a Bayesian method for estimating the state of a hidden Markov model (HMM). For instance, let k denote a time-step and consider the Markov chain (θk )k=1,2,3,... that takes value in the finite state-space S = {s1 , s2 , ..., sm }. There is an m × m matrix Λ of transition probabilities such that from time k to k + 1 there is the following probability of the event {θk+1 = si } given {θk = sj }, P(θk+1 = si |θk = sj ) = Λji Now suppose that this Markov chain is unobserved or hidden, but there is an observed process Yk that is a nonlinear function of θk , Yk = h(θk ) + “iid noise” A nonlinear filtering algorithm uses Bayes rule and the Markov property to formulate the posterior distribution of θk given Y1:k = (Y1 , Y2 , ...., Yk ) in a recursive manner: πki := P(θk = si |Y1:k ) P(θk = si , Yk |Y1:k−1 ) P(Yk |θk = si , Y1:k−1 )P(θk = si |Y1:k−1 ) = = P(Yk |Y1:k−1 ) P(Yk |Y1:k−1 ) P P(Yk |θk = si )P(θk = si |Y1:k−1 ) P(Yk |θk = si ) j P(θk = si , θk−1 = sj |Y1:k−1 ) = = P(Yk |Y1:k−1 ) P(Yk |Y1:k−1 ) 1 P P(Yk |θk = si ) j P(θk = si |θk−1 = sj , Y1:k−1 )P(θk−1 = sj |Y1:k−1 ) = P(Yk |Y1:k−1 ) P P(Yk |θk = si ) j P(θk = si |θk−1 = sj )P(θk−1 = sj |Y1:k−1 ) = P(Yk |Y1:k−1 ) 1 X j = P(Yk |θk = si ) Λji πk−1 ck j where ck is a normalizing constant. This simple probabilistic recursion is the main idea behind all filtering algorithms. Hidden Markov modeling and filtering algorithms have proven to be extremely useful in a number of applied areas including speech recognition, automatic target recognition, computer vision, and molecular biology. This thesis does considerable work with applications to human-tracking problems in computer vision, and also con- siders applications to finance which is an area where filtering is relatively undeveloped. In applications, the methods used to implement the associated algorithms are often non-trivial. In fact, in many cases it is nearly impossible to implement an algorithm whose output is at all close to the posterior distribution. In the past, approximations such as the extended Kalman Filter were used, and more recently as computer technology has improved it has become more practical to use Monte Carlo methods such as the Particle filter. One of the main topics of this thesis is fast-filtering algorithms for HMMs with multiple time-scales, where it is shown by example how to construct a filter of reduced-dimension when it is not necessary to compute the distribution for part of the state associated with a fast time-scale. Such filters are considerably faster because there are fewer dimensions that need to be integrated over in updating the posterior distributions. The most famous results in filtering theory are the Kalman Filter for linear Gauss- ian systems, and its counter-part from from SDE theory called the Kalman-Bucy fil- ter. For general nonlinear systems, there is the forward Baum-Welch equation which serves as the work-horse for all nonlinear filtering and Bayesian algorithms. Other 2 well-known nonlinear filtering results include the Kushner equation, the Zakai equa- tion, and the Wonham filter, as well as estimation methods such as the particle filter and Monte Carlo Markov Chain (MCMC). The work done in this thesis examines four separate nonlinear filtering problems in close detail. For each problem, a model is derived, a filtering algorithm is formulated, and numerical experiments are carried out. Chapters 2, 3, 4, and 5 of this dissertation are self-contained works each of which focusses on a specific problem in the field of nonlinear filtering with its own unique contribution. Chapter 2 is a novel application of filtering theory to finance and stochastic volatil- ity, but also serves as primer for the basic tools of nonlinear filtering such as the for- ward Baum-Welch equation, the backward Baum-Welch equation, and the parameter re-estimation procedure given by the expectation-maximization (EM) algorithm. His- torically, volatility is negatively correlated with the markets, and therefore is useful for hedging in times of uncertainty or possible downturn. One type of such hedging instruments are volatility swaps, whose returns rely on a rather crude time-series es- timate of the volatility process. As an alternative instrument, swap contracts can be written for the times series of filtering estimates, and the variance of returns on these contracts will be lower than the those of a traditional swap. The main results is that there is parity as swaps that use filtering volatility exhibit a mean/variance trade-off resulting in lower expected returns. Chapter 3 discusses a problem related to the identification/prediction of behav- ioral characteristics of an intelligent agent based on noisy video footage. The problem is motivated by the growing needs in surveillance for software that is capable of track- ing human beings. The software is developed with the goal of tracking an array of the agent’s attributes including patterns of motion (e.g. pose and position ), intentions, mood swings etc. Some of the attributes, (e.g. intensions) are not directly observable, and need to be inferred from the other attributes. Others, such as pose and location, are partially observable but the observations are corrupted by noise. 3 The approach to identification is based on nonlinear filtering-type algorithms and optimal change-point detection for partially observable double Markov chains (DMCs). Generally speaking, nonlinear filtering for partially observable DMCs is a particular case of the HMM filtering algorithm for vector-valued Markov chains. However, one of the main obstacles to efficient performance of this filtering algorithm is the curse of dimensionality. To some degree, this problem could be addressed by the introduction of particle filters, Rao-Blackwellization, and other methods, but the high dimensionality still remains a serious problem. Introduction of DMCs is just another step in the quest for reduction of computational complexity of HMMs. It compliments other methods (e.g. particle filters and Rao-Blackwellization) when possible. This ap- plication of DMC-based nonlinear filtering to analysis of video footage demonstrates the practical potential of the approach. Chapter 4 considers nonlinear filtering applications to target tracking based on a vector of multi-scaled models where some of the processes are rapidly mean reverting to their local equilibria. The algorithms are applicable to target tracking problems because multi-scaled models with fast mean-reversion (FMR) are a simple way to model latency in the response of tracking systems. A Kramers-Smoluchowski approx- imation is applied to the system of hidden states to obtain the main result, which is a simplified nonlinear filtering algorithms that is derived by exploiting the FMR struc- tures within the HMM. Specifically, a Baum-Welch recursion of reduced dimension is derived. The simplified algorithm is implemented in numerical simulations, and its efficiency and robustness is discussed. Chapter 5 is the most theoretical chapter in the thesis. The main part is the derivation of a fast nonlinear filtering algorithm of reduced dimension for a hidden Markov model with multiple time-scales. It is assumed that the motivation of the model and the nonlinear filtering algorithm are known, and then the bulk of the chapter focuses on the analytical/probabilistic details necessary to prove that the fast filter is optimal as the time-scale parameter increases to infinity. The main result 4 is a filter of reduced dimension for a particular class of multi-dimensional filtering problems where the hidden processes have multiple time scales. As possible future topics of research, each of the four problems in this thesis can be further analyzed to obtain new and improved results. It is without a doubt that filtering applications to finance will generate interest and that further data analysis will be necessary to strengthen the claims made thus far. The application to video tracking is a good start but to date there has been no serious attempt to construct filtering algorithms with parallel implementation in mind. It will be a considerable step forward for HMMs in computer vision if the rate at which posteriors are computed is significantly increased by implementing parallel versions of the nonlinear filtering algorithms. With regard to HMMs with multiple time-scales, there is a lot of analysis to be done for models with more parameters, as many of the proofs in this thesis will require extensions by the introduction of concepts such as moving-averages and heteroskedasticity. 5 CHAPTER 2 Filtering for Stochastic Volatility Models & Volatility Swap Contracts The Black-Scholes Theory says that volatility is positively correlated with the options market and is the key factor in pricing. Furthermore, since an option is an instrument whose values increases in times of uncertainty or possible downturn, volatility should be negatively correlated with the market. Indeed this trend is evi- dent in the data of volatility resources such as the VIX. However, volatility is not a process that is directly available for quotes, and so there is demand for things such as implied volatility, volatility surface estimates, and filtering volatility estimates. In addition, after an estimate of volatility has been computed there needs to be a correct interpretation of its meaning in the marketplace. With an understanding of volatility, an increased understanding of the brand of financial instruments known as volatility swaps will follow. Traditionally, a party enters into a volatility swap contract with the understanding that at specified terminal time they will receive the net difference between historical volatility and a specified strike price. For the S&P 500, such contracts are ‘fairly’ priced initially by the VIX which manipulates a portfolio of European call and put options to find the risk-neutral valuation of the contract’s expected payoff. Thus, the VIX is a form a implied volatility because it is computed from options data. Generally speaking, the returns on volatility swap contracts give the edge to the issuer of the contract, which can be empirically verified by the historically negative returns on volatility swaps. As an alternative to implied volatility, stochastic volatility and filtering algorithms deal directly with model parameters and observable data, and therefore return a form of historical volatility. Therefore, a time-series of filtering estimates for the volatility of the S&P 500 has fundamental differences from the well-regarded VIX, 6 but can provide a ‘better’ estimate of the average volatility over a time period than the traditional time-series estimate that is taken to be the sum of the squares on log-returns. Furthermore, upon examination of the data, one finds that the VIX and filtering estimates do have a lot in common, and in fact are a close fit if scaled properly. From this phenomenon of filtering tracking the VIX, it is clear that market forces are at work in a way that relates stochastic volatility to the VIX, and this needs to be interpreted in order to further understand the risk associated with the S&P 500 and other similar assets. In financial math, there has been a lot of work in stochastic and implied volatility, which are described in the book by Fouque, Papanicolaou and Sircar [19]. Information pertaining to the VIX such as its history, the dynamics of variance swaps, and data analysis, can be found in the paper Carr and Wu [12], and also in the paper by Cont and Kokholm [13]. Basic filtering theory is covered in the book by Jazwinski [25] and the paper by Rabiner [32] presents all the algorithms for filtering and parameter re- estimation. With specific regard to content of this chapter, the previous work done on this problem is relatively sparse, but the initiating of filtering for stochastic volatility is made in the paper by Cvitanic, Rozovsky and Zaliapan [15] where stochastic volatility is modeled with a Cox process. In this chapter, stochastic volatility modeling, implied volatility and nonlinear fil- tering come together in a way that is relatively new. A nonlinear filtering algorithm for a stochastic volatility model on a discrete space is formulated, along with asso- ciated algorithms for smoothing and parameter re-estimation. It is also shown that 30-day swap contracts offering a return on the average of the time-series of filtered volatility for S&P 500 volatility have returns with lower variance than the traditional swap contract, but the net returns have lower expectation. This chapter thoroughly analyzes volatility swaps for the S&P 500, but future work would needs to be done such as applying the same experiments to other indices such as the DOW, the Nasdaq, the S&P 100, and various stocks for which volatility 7 is traded. Another possibility would be to find a portfolio which replicates the time- series filtering estimates, but it is uncertain that such an instrument is attainable. 1. Historical Volatility, Variance-Swaps, and the VIX Let the financial markets be modeled in the probability space (Ω, F, P), where the price of a tradable asset is modeled as a stochastic processes St = the price of the stock or index at time t ≥ 0. A widely accepted stochastic volatility model for St is the following SDE (1) dSt = µt St dt + σt St dWt where µt is a process which models annualized growth, σt is a process which models annualized volatility, and Wt is an independent Brownian motion. Quotes on St are available throughout the trading day, but the continuum of prices is not observable. For some time T > 0, there is a discrete set of times (tk )N +1 k=0 at which quotes are available, Stk = the observed quote on the asset at time tk ∈ [0, T ] The solution of (1) on [tk , tk+1 ] is Z tk+1   Z tk+1  1 2 Stk+1 = Stk exp µτ − στ dτ + στ dWτ tk 2 tk which leads to a the following time-series model of log-returns: Z tk+1   Z tk+1 1 2 (2) Yk := log(Stk+1 /Stk ) = µτ − στ dτ + στ dWτ tk 2 tk For simplicity, assume that there is an increment ∆t such that the quotes are given at equidistant times, (3) 0 = t0 < t1 < ..... < tN +1 = T, where tk+1 − tk ≡ ∆t, ∀k 8 Define the average volatility on [0, T ] as 1 T 2 Z 2 (4) σ[0,T ] = σ dτ T 0 τ and consider the time-series average of the squared of returns N 1X 2 (5) RV[0,T ] := Y T k=0 k which converges in mean to (4) as the sample size goes to a continuum, " N Z # tk+1  Z T  1X 2 1 2 E[RV[0,T ] ] = E στ dτ + o(∆t) =E σ dτ + O(∆t) T k=0 tk T 0 τ 2 2 (6) = E[σ[0,T ] ] + O(∆t) → E[σ[0,T ] ] as N % ∞ The average in (5) is commonly referred to as realized or historical volatility, but in practice it is not possible to obtain a continuum of quotes in the interval [0, T ], and so historical volatility is only an estimate of the average volatility. However, historical volatility is often used in a brand of derivative contracts called variance swaps. Variance swaps are contracts that pay the net difference of histor- ical volatility and a pre-specified strike. For a strike-price Kvol , a long position in a swap on volatility pays the difference in historical volatility and the strike price, 2 variance-swap payoff = RV[0,T ] − Kvol 1.1. ‘Fair’ Valuation of Volatility Swaps. For a time t > 0, the price of the swap is dictated by the market. Initially, at time t = 0, by arbitrage arguments, the ‘fair’ price should be the discounted risk-neutral expectation of the expected payoff, fair swap price|t=0 = e−rT E∗0 [RV[0,T ] ] − Kvol 2 ≈ e−rT E∗0 [σ[0,T 2 2   ] ] − K vol where r > 0 is the risk-free rate of interest and E∗t [·] is the expectation under the risk-neutral measure conditioned on the market data to time t. The risk free measure is defined in terms of a Martingale, ( 2 ) 1 T µτ − r Z T dP∗ Z  µτ − r = MT = exp − dτ − dWτ dP 2 0 στ 0 στ 9 under which dSt = rSt dt + σt St dWt∗ where Wt∗ is P∗ -Brownian motion and dWt∗ = µt −r σt dt + dWt , and so the limit in (6) holds under the risk-neutral measure. However, it is not immediately clear how to obtain an explicit expression for this fair price. Fortunately, a clever portfolio can be constructed and arbitrate theory can be applied to obtain this price. Suppose that at time t = 0 there is the following forward spot rate for ST , F = spot rate for a futures contract in ST at time t = 0 and there is a risk free rate of return r such that F = E∗0 ST = S0 erT . For any strike price K, let C(T, t; K) and P (T, t; K) denote the risk-neutral arbitrage-free prices of a European call and put, respectively. Then, for any time t ∈ [0, T ] consider the following portfolios of out-of-the money calls and puts: portfolio 1 = {1 of each out-of-the-money European call option} portfolio 2 = {1 of each out-of-the-money European put option} Letting the prices of these two portfolios be given by Φ1t and Φ2t , respectively, Z ∞ Z F 1 r(T −t) C(T, t; K) 2 r(T −t) P (T, t; K) Φt = e 2 dK, Φt = e dK F K 0 K2 by arbitrage-free pricing theory, Φ1t can be rewritten as ∞ ∞ ∞ s s−K s−K Z Z Z Z r(T −t) −r(T −t) Φ1t =e e ρT −t (s)dsdK = ρT −t (s) dKds F K K2 F F K2 d ∗ where ρT −t (s) = ds P (ST ≤ s|St ). Similarly for Φ2t , we have F Z K Z F Z F K −s K −s Z 2 r(T −t) −r(T −t) Φt = e e 2 ρT −t (s)dsdK = ρT −t (s) dKds 0 0 K 0 s K2 Summing them together and integrating over K, there is the following expression for the joint portfolio for any time, ∞ F ∞   K −s s−F Z Z Z Φ1t + Φ2t = ρT −t (s) dKds = ρT −t (s) log(F/s) + ds 0 s K2 0 F 10 E∗t ST − F = E∗t [log(F/ST )] + F Therefore, a short position in the log-returns on a futures contract can be hedged with a short position in a futures contract and a long position in Φ1T and Φ2T , ST − F − log(ST /F ) = − + Φ1T + Φ2T F In particular, at time t = 0, the model for the index price says that ST = RT (r−.5σt2 )dt+ 0T σt dWt∗ R S0 e 0 where Wt∗ is Brownian motion under the risk-neutral mea- sure, as so the expected logarithm of returns can be written as the integral of the volatility, Z T h  RT .5σt2 dt+ RT σt dWt∗ i 1 Φ10 + Φ20 = −E∗0 log(ST /F ) = −E∗0 log e − 0 0 = E∗0 [σt2 ]dt 2 0 where E∗ [·] denotes the risk neutral expectation. Thus there is the following fair valuation of the average volatility on [0, T ], 2 1 E∗0 σ[0,T 2 ] = (Φ + Φ20 ) T 0 and the fair price of the volatility swap with strike Kvol is given in terms of these portfolios: 2 1 fair price of swap|t=0 = (Φ + Φ20 ) − Kvol 2 T 0 2 1.2. The VIX. The price T (Φ10 + Φ20 ) with T = 30 days is what the VIX ap- proximates by using all options data available on the S&P 500 with T -many days to expiry. In computing the VIX, the continuum of strike prices is approximated with the set of strike prices (Ki )i=1,2,3,.... provided by the market, ! 2erT X C(tk + T, tk ; Ki ) X P (tk + T, tk ; Ki ) (7) V IXt2k = ∆K i + ∆Ki T K >F Ki2 K