程式師世界 >> 編程語言 >> 網頁編程 >> PHP編程 >> 關於PHP編程 >> 使用PHP導出Word文檔的原理和實例

使用PHP導出Word文檔的原理和實例

編輯：關於PHP編程

原理

一般，有2種方法可以導出doc文檔，一種是使用com，並且作為php的一個擴展庫安裝到服務器上，然後創建一個com，調用它的方法。安裝過office的服務器可以調用一個叫word.application的com，可以生成word文檔，不過這種方式我不推薦，因為執行效率比較低（我測試了一下，在執行代碼的時候，服務器會真的去打開一個word客戶端）。理想的com應該是沒有界面的，在後台進行數據轉換，這樣效果會比較好，但是這些擴展一般需要收費。
第2種方法，就是用PHP將我們的doc文檔內容直接寫入一個後綴為doc的文件中即可。使用這種方法不需要依賴第三方擴展，而且執行效率較高。
word本身的功能還是很強大的，它可以打開html格式的文件，並且能夠保留格式，即使後綴為doc，它也能識別正常打開。這就為我們提供了方便。但是有一個問題，html格式的文件中的圖片只有一個地址，真正的圖片是保存在其他地方的，也就是說，如果將HTML格式寫入doc中，那麼doc中將不能包含圖片。那我們如何創建包含圖片的doc文檔呢？我們可以使用和html很接近的mht格式。
mht格式和html很類似，只不過在mht格式中，外部鏈接進來的文件，比如圖片、Javascript、CSS會被base64進行編碼存儲。因此，單個mht文件就可以保存一個網頁中的所有資源，當然，相比html，它的尺寸也會比較大。
mht格式能被word識別嗎？我將一個網頁保存成mht，然後修改後綴名為doc，再用word打開，OK，word也可以識別mht文件，並且可以顯示圖片。
好了，既然doc可以識別mht，下面就是考慮如何將圖片放入mht了。由於html代碼中的圖片的地址都是寫在img標簽的src屬性中，因此，只要提取html代碼中的src屬性值，就可以獲得圖片地址。當然，有可能您獲取到的是相對路徑，沒關系，加上URL的前綴，改成絕對路徑就可以了。有了圖片地址，我們就可以通過file_get_content函數獲取到圖片文件的具體內容，然後調用base64_encode函數將文件內容編碼成base64編碼，最後插入到mht文件的合適位置即可。
最後，我們有兩種方法將文件發送給客戶端，一種是先在服務器端生成一個doc文檔，然後將這個doc文檔的地址記錄下來，最後，通過header("location:xx.doc");就可以讓客戶端下載這個doc。還有一種是直接發送html請求，修改HTML協議的header部分，將它的content-type設置為application/doc，將content-disposition設置為attachment，後面跟上文件名，發送完html協議以後，直接將文件內容發送給客戶端，也可以讓客戶端下載到這個doc文檔。

實現

通過以上的原理介紹，相信大家應該對實現的過程有個初步的了解了，下面我給出一個導出函數，這個函數可以將HTML代碼導出成一個mht文檔，參數有3個，其中後2個為可選參數
content:要轉換的HTML代碼
absolutePath: 如果HTML代碼中的圖片地址都是相對路徑，那麼這個參數就是HTML代碼中缺少的絕對路徑。
isEraseLink:是否去掉HTML代碼中的超鏈接
返回值為mht的文件內容，您可以通過file_put_content將它保存成後綴名為doc的文件
這個函數的主要功能其實就是分析HTML代碼中的所有圖片地址，並且依次下載下來。獲取到了圖片的內容以後，調用MhtFileMaker類，將圖片添加到mht文件中。具體的添加細節，封裝在MhtFileMaker類中了。

復制代碼代碼如下:
/**
* 根據HTML代碼獲取word文檔內容
* 創建一個本質為mht的文檔，該函數會分析文件內容並從遠程下載頁面中的圖片資源
* 該函數依賴於類MhtFileMaker
* 該函數會分析img標簽，提取src的屬性值。但是，src的屬性值必須被引號包圍，否則不能提取
*
* @param string $content HTML內容
* @param string $absolutePath 網頁的絕對路徑。如果HTML內容裡的圖片路徑為相對路徑，那麼就需要填寫這個參數，來讓該函數自動填補成絕對路徑。這個參數最後需要以/結束
* @param bool $isEraseLink 是否去掉HTML內容中的鏈接
* by www.jb51.net
*/
function getWordDocument( $content , $absolutePath = "" , $isEraseLink = true )
{
    $mht = new MhtFileMaker();
    if ($isEraseLink)
        $content = preg_replace('/<a\s*.*?\s*>(\s*.*?\s*)<\/a>/i' , '$1' , $content);   //去掉鏈接

    $images = array();
    $files = array();
    $matches = array();
    //這個算法要求src後的屬性值必須使用引號括起來
    if ( preg_match_all('/<img[.\n]*?src\s*?=\s*?[\"\'](.*?)[\"\'](.*?)\/>/i',$content ,$matches ) )
    {
        $arrPath = $matches[1];
        for ( $i=0;$i<count($arrPath);$i++)
        {
            $path = $arrPath[$i];
            $imgPath = trim( $path );
            if ( $imgPath != "" )
            {
                $files[] = $imgPath;
                if( substr($imgPath,0,7) == 'http://')
                {
                    //絕對鏈接，不加前綴
                }
                else
                {
                    $imgPath = $absolutePath.$imgPath;
                }
                $images[] = $imgPath;
            }
        }
    }
    $mht->AddContents("tmp.html",$mht->GetMimeType("tmp.html"),$content);

    for ( $i=0;$i<count($images);$i++)
    {
        $image = $images[$i];
        if ( @fopen($image , 'r') )
        {
            $imgcontent = @file_get_contents( $image );
            if ( $content )
                $mht->AddContents($files[$i],$mht->GetMimeType($image),$imgcontent);
        }
        else
        {
            echo "file:".$image." not exist!<br />";
        }
    }

    return $mht->GetFile();
}

使用方法：

復制代碼代碼如下:
$fileContent = getWordDocument($content,"http://www.jb51.net/Music/etc/");
$fp = fopen("test.doc", 'w');
fwrite($fp, $fileContent);
fclose($fp);

其中，$content變量應該是HTML源代碼，後面的鏈接應該是能填補HTML代碼中圖片相對路徑的URL地址
注意，在使用這個函數之前，您需要先包含類MhtFileMaker，這個類可以幫助我們生成Mht文檔。

復制代碼代碼如下:
<?php
/***********************************************************************
Class:        Mht File Maker
Version:      1.2 beta
Author:       Wudi <[email protected]>
Description: The class can make .mht file.
***********************************************************************/

class MhtFileMaker{
    var $config = array();
    var $headers = array();
    var $headers_exists = array();
    var $files = array();
    var $boundary;
    var $dir_base;
    var $page_first;

function MhtFile($config = array()){

}

    function SetHeader($header){
        $this->headers[] = $header;
        $key = strtolower(substr($header, 0, strpos($header, ':')));
        $this->headers_exists[$key] = TRUE;
    }

    function SetFrom($from){
        $this->SetHeader("From: $from");
    }

    function SetSubject($subject){
        $this->SetHeader("Subject: $subject");
    }

    function SetBoundary($boundary = NULL){
        if ($boundary == NULL) {
            $this->boundary = '--' . strtoupper(md5(mt_rand())) . '_MULTIPART_MIXED';
        } else {
            $this->boundary = $boundary;
        }
    }

    function SetBaseDir($dir){
        $this->dir_base = str_replace("\\", "/", realpath($dir));
    }

    function SetFirstPage($filename){
        $this->page_first = str_replace("\\", "/", realpath("{$this->dir_base}/$filename"));
    }

    function AutoAddFiles(){
        if (!isset($this->page_first)) {
            exit ('Not set the first page.');
        }
        $filepath = str_replace($this->dir_base, '', $this->page_first);
        $filepath = 'http://mhtfile' . $filepath;
        $this->AddFile($this->page_first, $filepath, NULL);
        $this->AddDir($this->dir_base);
    }

    function AddFile($filename, $filepath = NULL, $encoding = NULL){
        if ($filepath == NULL) {
            $filepath = $filename;
        }
        $mimetype = $this->GetMimeType($filename);
        $filecont = file_get_contents($filename);
        $this->AddContents($filepath, $mimetype, $filecont, $encoding);
    }

    function AddContents($filepath, $mimetype, $filecont, $encoding = NULL){
        if ($encoding == NULL) {
            $filecont = chunk_split(base64_encode($filecont), 76);
            $encoding = 'base64';
        }
        $this->files[] = array('filepath' => $filepath,
                               'mimetype' => $mimetype,
                               'filecont' => $filecont,
                               'encoding' => $encoding);
    }

    function CheckHeaders(){
        if (!array_key_exists('date', $this->headers_exists)) {
            $this->SetDate(NULL, TRUE);
        }
        if ($this->boundary == NULL) {
            $this->SetBoundary();
        }
    }

    function CheckFiles(){
        if (count($this->files) == 0) {
            return FALSE;
        } else {
            return TRUE;
        }
    }

    function MakeFile($filename){
        $contents = $this->GetFile();
        $fp = fopen($filename, 'w');
        fwrite($fp, $contents);
        fclose($fp);
    }

    function GetMimeType($filename){
        $pathinfo = pathinfo($filename);
        switch ($pathinfo['extension']) {
            case 'htm': $mimetype = 'text/html'; break;
            case 'html': $mimetype = 'text/html'; break;
            case 'txt': $mimetype = 'text/plain'; break;
            case 'cgi': $mimetype = 'text/plain'; break;
            case 'php': $mimetype = 'text/plain'; break;
            case 'css': $mimetype = 'text/css'; break;
            case 'jpg': $mimetype = 'image/jpeg'; break;
            case 'jpeg': $mimetype = 'image/jpeg'; break;
            case 'jpe': $mimetype = 'image/jpeg'; break;
            case 'gif': $mimetype = 'image/gif'; break;
            case 'png': $mimetype = 'image/png'; break;
            default: $mimetype = 'application/octet-stream'; break;
        }
        return $mimetype;
    }
}
?>

上面討論了通過mht文件，來實現PHP導出doc格式的。這種方法可以解決一個難題，就是使導出的doc文件中包含圖片，當然，如果您要包含更多的內容，比如CSS樣式表，只需要用正則表達式分析HTML代碼中的link標簽，提取css樣式文件的地址，然後讀取並編碼成base64，最後加入到mht文件中就可以了。