# Retrieving the source code from a webpage?

**URL:** <https://forum.kirupa.com/t/retrieving-the-source-code-from-a-webpage/116837>\
**Category:** Uncategorized\
**Created:** [July 21, 2004, 7:30pm UTC](https://forum.kirupa.com/t/retrieving-the-source-code-from-a-webpage/116837 "2004-07-21T19:30:34Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![RoboLlama](https://avatars.discourse-cdn.com/v4/letter/r/bc79bd/32.png) [@RoboLlama](https://forum.kirupa.com/u/RoboLlama)\
**Post date:** [July 21, 2004, 7:30pm UTC](https://forum.kirupa.com/t/retrieving-the-source-code-from-a-webpage/116837/1 "2004-07-21T19:30:34Z")

</div>

Is there any way I can use Flash MX 2004 to retrieve the source code from a webpage? For example, I give Flash the string “[http://www.google.com](http://www.google.com)”, and Flash gives me the source code for google’s home page.

If not, does anyone know what programming language I can use to do this?

---

<div class="post-metadata">

**Author:** ![system](https://canada1.discourse-cdn.com/flex011/uploads/kirupa/original/3X/6/2/621e5c11736f46532526e61be85940af4230f3e5.png) [@system](https://forum.kirupa.com/u/system)\
**Post date:** [July 21, 2004, 9:52pm UTC](https://forum.kirupa.com/t/retrieving-the-source-code-from-a-webpage/116837/2 "2004-07-21T21:52:38Z")

</div>

No. You can’t view the server side coding on webpages.

---

<div class="post-metadata">

**Author:** ![system](https://canada1.discourse-cdn.com/flex011/uploads/kirupa/original/3X/6/2/621e5c11736f46532526e61be85940af4230f3e5.png) [@system](https://forum.kirupa.com/u/system)\
**Post date:** [July 21, 2004, 9:54pm UTC](https://forum.kirupa.com/t/retrieving-the-source-code-from-a-webpage/116837/3 "2004-07-21T21:54:08Z")

</div>

but you can see the finished source code, there’s a php function that’ll look up source code for pages, search for it on their site 😉

---

<div class="post-metadata">

**Author:** ![system](https://canada1.discourse-cdn.com/flex011/uploads/kirupa/original/3X/6/2/621e5c11736f46532526e61be85940af4230f3e5.png) [@system](https://forum.kirupa.com/u/system)\
**Post date:** [July 21, 2004, 10:01pm UTC](https://forum.kirupa.com/t/retrieving-the-source-code-from-a-webpage/116837/4 "2004-07-21T22:01:38Z")

</div>

something like

> webcode=new LoadVars()  
> webcode.load(“[http://www.google.de](http://www.google.de)”)  
> \_root.onEnterFrame=function(){  
> trace(webcode)  
> }  
> ?

---

<div class="post-metadata">

**Author:** ![system](https://canada1.discourse-cdn.com/flex011/uploads/kirupa/original/3X/6/2/621e5c11736f46532526e61be85940af4230f3e5.png) [@system](https://forum.kirupa.com/u/system)\
**Post date:** [July 21, 2004, 10:07pm UTC](https://forum.kirupa.com/t/retrieving-the-source-code-from-a-webpage/116837/5 "2004-07-21T22:07:28Z")

</div>

I thought you were kidding :P, that actually kind of works, other than it traces like a million times, but there seems to be html in there, I saw cellspacing?

---

<div class="post-metadata">

**Author:** ![system](https://canada1.discourse-cdn.com/flex011/uploads/kirupa/original/3X/6/2/621e5c11736f46532526e61be85940af4230f3e5.png) [@system](https://forum.kirupa.com/u/system)\
**Post date:** [July 21, 2004, 10:33pm UTC](https://forum.kirupa.com/t/retrieving-the-source-code-from-a-webpage/116837/6 "2004-07-21T22:33:46Z")

</div>

I don’t know if this would help, but when you type in ‘view-source:[http://www.google.com](http://www.google.com)’ minus the quotes (on IE at least), it pops up the source code. Don’t know if this might help in flash…

[EDIT]  
Sticking

```auto
getURL ("view-source:http://www.google.com");

```

in the first keyframe of a blank movie seems to work (pops up the source code in Notepad for me), maybe you could make some use of this after all?  
[/EDIT]

Hope it helps! 🍺

---

<div class="post-metadata">

**Author:** ![system](https://canada1.discourse-cdn.com/flex011/uploads/kirupa/original/3X/6/2/621e5c11736f46532526e61be85940af4230f3e5.png) [@system](https://forum.kirupa.com/u/system)\
**Post date:** [July 22, 2004, 12:10am UTC](https://forum.kirupa.com/t/retrieving-the-source-code-from-a-webpage/116837/7 "2004-07-22T00:10:12Z")

</div>

Thank you McGiver, that was exactly what I was looking for.

By the way, I modified the code, so that it only traces once, and so that it decodes all the escape sequences.

[AS]webcode = new LoadVars()  
webcode.load(“[http://www.google.com](http://www.google.com)”)  
webcode.onLoad = function()  
{  
trace(unescape(this))  
}[/AS]

This traces a readable source code to the webpage.

Thanks again!

---

<div class="post-metadata">

**Author:** ![system](https://canada1.discourse-cdn.com/flex011/uploads/kirupa/original/3X/6/2/621e5c11736f46532526e61be85940af4230f3e5.png) [@system](https://forum.kirupa.com/u/system)\
**Post date:** [July 22, 2004, 1:22am UTC](https://forum.kirupa.com/t/retrieving-the-source-code-from-a-webpage/116837/8 "2004-07-22T01:22:07Z")

</div>

Wow - this is really neat! I had no idea that could be done. Would either of you mind if I write a small article about this on the site? I’ll definitely give both of you credit for asking and answering 🙂

---

<div class="post-metadata">

**Author:** ![system](https://canada1.discourse-cdn.com/flex011/uploads/kirupa/original/3X/6/2/621e5c11736f46532526e61be85940af4230f3e5.png) [@system](https://forum.kirupa.com/u/system)\
**Post date:** [July 22, 2004, 1:29am UTC](https://forum.kirupa.com/t/retrieving-the-source-code-from-a-webpage/116837/9 "2004-07-22T01:29:05Z")

</div>

hmm, concerning the article:  
I just looked over it again, and it produces some problems:

-on the end of every code is a “&onLoad=%5Btype%20Function%5D” (== “&onLoad=[type Function]”). =\>easy to solve, no problem at all.

-code is full of %20 and stuff. =\>no problem at all, since you can decode it easyly to its original form ("%20"==" ")

- stuff like "&nbsp;" (another way to display " " == coded space) causes problems (to be precise the “&” causes the problems) . I guess this happens because the “&” is the standart serverstring delimiter (i.e. &aaa=hallo&bbb=peops are stored as myloadvars.aaa=hallo,… and not given back in the correct order)  
so if you “read” a source (i.e. google) the code is mixed up and some parts (i.e. "&nbsp;\<a href=“abc”\>hallo\</a\>&nbsp;\<a href=“abc”\>at all\</a\> replace each other because the "&nbsp;\<a href=" is seen as a variable declaration) are even missing

here the improved code

```auto
String.prototype.tochars = function() {
   	var current = this;
   	var final;
   	var indpos;
   	var indpos2=this.indexOf("&onLoad=[type Function]",0)
   	this=this.substr(0,indpos2)+this.substr(indpos2+23)
   	indpos = this.indexOf("%", 0);
   	while (indpos != -1) {
 		this = this.substr(0, indpos)+String.fromCharCode(Number("0x"+this.charAt(indpos+1)+this.charAt(indpos+2)))+this.substr(indpos+3);
   		indpos = this.indexOf("%", 0);
   	}
   	return (this);
   };
   webcode = new LoadVars();
   webcode.load("http://www.google.com/");
   webcode.onLoad = function() {
   	strangestring = webcode.toString();
   	_root.createTextField("showhtml", 10, 5, 5, 890, 590);
   	htmlstring = strangestring.tochars();
   	_root.showhtml.text = htmlstring.tochars();
   };

```

---

<div class="post-metadata">

**Author:** ![system](https://canada1.discourse-cdn.com/flex011/uploads/kirupa/original/3X/6/2/621e5c11736f46532526e61be85940af4230f3e5.png) [@system](https://forum.kirupa.com/u/system)\
**Post date:** [July 22, 2004, 6:21am UTC](https://forum.kirupa.com/t/retrieving-the-source-code-from-a-webpage/116837/10 "2004-07-22T06:21:03Z")

</div>

Cool, I wouldn’t mind if you wrote an article Kirupa.

McGiver: you could do that, or you could just use the built in unescape function (see the code in my other post in this thread); It does all that stuff for you.

And a suggestion for the article: it should talk about the Flash security issue. That is, how this technique works fine until you try to put it into a webpage or upload it onto the internet. Maybe it should include how to write a server-side proxy in order to fix this problem?

Following is Macromedia’s documented ASP proxy. I’ve used it, and it works well.

```auto
<%@ LANGUAGE=VBScript%>
<% 
	Response.Buffer=True 
	Dim MyConnection, TheURL
	
	' Specifying the URL
	TheURL = "http://www.macromedia.com/desdev/resources/macromedia_resources.xml"
	
	 Set MyConnection = Server.CreateObject("Microsoft.XMLHTTP")
	' Connecting to the URL
	MyConnection.Open "GET", TheURL, False
		
	' Sending and getting data
	MyConnection.Send 
	TheData = MyConnection.responseText

	'Set the appropriate content type
	Response.ContentType = MyConnection.getResponseHeader("Content-Type")
	Response.Write (TheData)

	Set MyConnection = Nothing
%>

```

---

<div class="post-metadata">

**Author:** ![system](https://canada1.discourse-cdn.com/flex011/uploads/kirupa/original/3X/6/2/621e5c11736f46532526e61be85940af4230f3e5.png) [@system](https://forum.kirupa.com/u/system)\
**Post date:** [July 22, 2004, 4:56pm UTC](https://forum.kirupa.com/t/retrieving-the-source-code-from-a-webpage/116837/11 "2004-07-22T16:56:33Z")

</div>

I wrote something that returns the source code without the "&"s

```auto
<?php
 $file=fopen($url,"rb");
 $content=fread($file,102400);
 echo str_replace("&", "%26", $content);
 fclose($file)
 ?>

```

```auto
webcode = new LoadVars();
 sendvar = new LoadVars();
 sendvar.url="http://www.google.com/"
 sendvar.sendAndLoad("read.php",webcode);
 webcode.onLoad = function() {
 	strangestring = webcode.toString();
 	_root.createTextField("showhtml", 10, 5, 5, 890, 590);
 	htmlstring = unescape(strangestring);
 	var indpos = htmlstring.indexOf("&onLoad=[type Function]", 0);
 	htmlstring = htmlstring.substr(0, indpos)+htmlstring.substr(indpos+23);
 	_root.showhtml.text = htmlstring;
 };
 

```
